How accurate is it, on documents we did not choose
Precision, recall and F1 for KeptPDF redaction and automatic document splitting, corpus by corpus, with the method written down and the poor scores printed next to the good ones. Splitting measured 20 September 2026, redaction 3 September 2026.
This page publishes what KeptPDF actually scores on named corpora of real public records, with the scoring method and the corpus sizes attached. The honest one-line summary: redaction is strong and splitting is not. Redaction reads F1 82.5 to 92.5 on six of seven hand-labelled corpora. Automatic document splitting reads 96.9 on the discovery productions its rules were tuned against and 55.6 on a stack of city permit forms.
We publish the low numbers because a number with no corpus behind it is not a number. If you are evaluating tools, the useful question is not “how accurate is it” but “on what, measured how, and where does it fall over”. Those three answers are below.
How we score
Recall is how much of what should have been found was found. Precision is how much of what was found should have been. F1 squeezes the two into one number, so a tool cannot win by flagging everything or by flagging nothing.
A worked example. A page holds 10 pieces of personal data. The tool marks 8 of them and also marks 2 things that are not personal data. Recall is 8 of 10, so 80%. Precision is 8 of the 10 marks, so 80%. F1 is 80.
Redaction is scored twice, because the product has two tiers. R-auto is what the automatic pass marks on its own. R+rev also counts the items the tool holds back and puts in front of a person to confirm. The gap between the two columns is the work the reviewer is doing, and on a public records production that reviewer is not optional.
For splitting, a cut is the tool saying “a new document starts here”. Precision is how many of its cuts were real starts. Recall is how many of the real starts it found. A 48-page property file holding 9 filings has 8 real starts inside it, among 47 page gaps where a cut could be proposed. We report F1 on cuts and never plain accuracy, because starts are rare: they are about 9% of the page gaps in the discovery corpus, so a tool that proposes nothing at all would score over 90% accuracy while finding nothing.
The keys are hand made, and hand made keys can be wrong. Where a document already has a publisher index or a filed page count, we use it. Where it does not, a person read the pages. Twice now the key was the thing at fault rather than the engine, and both were corrected rather than left standing: an earlier labelling pass on the medical slice was scoring clinical dates as false alarms, and completing the labels to the standard the engine implements moved date precision from 26.1 to 71.3 with no code change at all.
One more limit worth stating before the tables. The redaction figures come from a harness that labels the gold from the page’s extracted text. It measures the detectors, not the whole pipeline, so it is blind by construction to a scanned page with no text layer and to anything that exists only as pixels: a signature, an inked-in form box, a stamp. Those are audited separately against labelled page images, and that audit is not what this table shows.
| Corpus | What it is | Pages | Items | P | R-auto | R+rev | F1 |
|---|---|---|---|---|---|---|---|
| trials | Clinical trial packets filed on ClinicalTrials.gov | 263 | 262 | 91.6 | 93.5 | 95.4 | 92.5 |
| productions | Real merged discovery productions somebody filed as one PDF | 130 | 270 | 93.7 | 90.7 | 97.8 | 92.2 |
| realfiles | Federal Register issues as the publisher issued them | 120 | 325 | 90.6 | 93.2 | 97.5 | 91.9 |
| collected | Public-domain government, court, medical and academic PDFs | 173 | 529 | 90.3 | 87.9 | 93.4 | 89.1 |
| matters | Court matter packets: one matter, one producer, one subject | 318 | 501 | 86.5 | 90.8 | 96.2 | 88.6 |
| hr | Employment and personnel records from court productions | 317 | 761 | 90.6 | 75.8 | 91.3 | 82.5 |
| municipal | City hall paper: property files, council matters, agenda packets | 40 | 173 | 78.9 | 30.6 | 72.8 | 44.1 |
That is 1,361 hand-labelled pages carrying 2,821 items of personal data. Two more corpora exist and are deliberately blank: finance and medical carry no labelled text gold in the current board, so they are not scored and no number is claimed for them. They will appear here when they are labelled, whatever they say.
Municipal is the worst row on the board and it is also the row that matters most to a city clerk, so read it carefully. Automatic recall is 30.6, recall including the review tier is 72.8. Most of that gap is street addresses on city forms, where the address of the office that issued the paper sits on the same page as the address of the resident. The tool finds them and holds them for a person to confirm rather than blacking them out unprompted, because getting that pair the wrong way round without asking would be worse than asking.
| Corpus | What it is | Size | Cut F1 |
|---|---|---|---|
| productions | 13 real merged productions, seams read by a person | 1,293 pages, 111 documents | 96.9 |
| realfiles | 12 Federal Register issues, keyed from the publisher index | 4,218 pages, 1,044 documents | 94.8 |
| trials | 27 clinical trial packets, keyed by the filed page counts | 4,148 pages, 96 documents | 65.4 |
| matters | 73 matter packets across five genres | 6,935 pages, 401 documents | 58.8 |
| IQM2 agenda packets | 16 council agenda packets, keyed by the packet outline | 969 pages | 64.9 |
| municipal, all slices | Laserfiche property files plus Legistar council matters | 1,240 pages, 197 documents | 52.3 |
| municipal property files | 24 Laserfiche property files on their own | 407 pages, 136 documents | 55.6 |
The property file row breaks down as precision 64.0 and recall 49.1: 55 cuts right, 31 wrong, 57 seams missed. That is the number to plan around if your backfile is a stack of one-page filings on the same city form.
Two of these rows flatter us and we would rather say so than have you find out. The productions corpus is the set the splitting rules were tuned against, so 96.9 is an in-sample score and must not be read as accuracy on documents nobody has seen. The Federal Register slice is genuinely real but genuinely narrow: it is one publisher, twelve issues, and most of its documents begin partway down a sheet the previous document is still using, which a page-level cut cannot reach at all. The city hall slices are the harder and more recent test, and they are the ones that read in the fifties.
The image-only scan rehearsal
A backfile usually arrives as one scanned PDF with no text layer. No corpus we own started that way, so we built the rehearsal: municipal property files rendered back to greyscale pages at scanner resolution and stapled into one image-only PDF, with a key recording the true first page of every filing.
On the three biggest rehearsal packets, 145 pages in total, the command scored cuts precision 41.9, recall 54.2, F1 47.3, and sorted the documents it did find into the right pile 15 times out of 16. It ran at about 2.16 seconds a page end to end on a 16-core machine with 30.8 GB of RAM, reading every page before scoring anything. Measured 2026-09-05.
Naming is better than cutting. On the municipal corpus the reader that decides what each document is (permit, application, certificate, recorded instrument, correspondence) was right on 125 of 136 documents, 91.9%, when it was allowed to read the scanned pages, against 110 of 136, 80.9%, from the text layer alone.
What we get wrong
- A stack of identical forms reads as one document. A Laserfiche property file is a run of one-page filings on the same city form, so a tool that works from what is on the page has very little to tell it that a new filing has started. It is still the hardest shape in the splitting table, and it is why the property row sits at 55.6.
- Hand-lettered oversize sheets name themselves as nothing. Plats, site plans and survey sheets carry no printed title a reader can recover. The plans pile scores 0 and the survey pile 40% recall, and that does not move until something other than the words on the sheet can answer.
- Municipal redaction leans on the reviewer. Automatic recall of 30.6 against 72.8 with review means a person is doing most of the work on city hall paper today. We would rather report that split than average it away.
- Long single documents draw false cuts. A cover page, a table page and a narrative page inside one municipal budget genuinely do not resemble each other, and that is where the wrong cuts concentrate.
- Two corpora are unscored. Finance and medical have no labelled text gold. An unscored corpus is a hole in the board, not a pass.
- One key reads low by definition. In the Laserfiche and Legistar keys, one filed permit is one document even when the city scanned an application, an inspection card and two signed approvals into it. The tool cutting finer than that is still scored as wrong, so property file precision is a floor rather than a fair reading.
How to reproduce this
Every corpus here is public records or public-domain publications. The PDFs are not redistributed; each slice has a fetch script that pulls the files back from their original public sources, and the committed record is the key plus a manifest carrying the URL, the byte size and the page count of every file.
- Splitting, member page count keys:
node eval/split/matter-eval.js --corpus <dir> - Splitting, hand-read and outline keys:
node eval/split/prod-eval.js --corpus <dir> - Image-only rehearsal:
node eval/split/make-scan-packet.jsthennode eval/split/scan-packet-eval.js --in <dir> - Redaction, text-labelled gold:
node eval/redact/text-gold.js <corpus> - Redaction, image-labelled gold:
node eval/redact/vision-gold.js - Redaction, synthetic regression set:
node eval/redact/run-eval.js. This is a generated 60-document guard that catches regressions, not a real-file result, and its score is not on this page.
The point of publishing the harness rather than only the score is that you can point it at documents we have never seen. If you run it on your own repository and get a worse number than the table, we would like to know, and the misses are more useful to us than the hits.
How to read a vendor accuracy claim
Most of the redaction vendor pages we checked publish a time-saving percentage or an “AI-powered” claim and no corpus, no metric and no method. A claim with no corpus attached cannot be checked, cannot be reproduced, and cannot be compared with anything. Four questions cut through it quickly.
- On what documents? Named and fetchable, or unnamed. If a vendor cannot say which files, the number describes nothing.
- Measured how? Precision and recall separately, or one blended figure. Ask for both, because a tool can buy recall with false positives all day.
- Who labelled it, and was the key ever wrong? A benchmark that has never corrected its own key has probably never been examined hard.
- Where does it fail? A published benchmark with no bad rows in it is a marketing table.
Two things worth crediting. On the splitting side, Extend.ai publishes a real benchmark with a released dataset and modest numbers, which is the template we followed here. On the redaction side the measured work we could find is academic rather than commercial, and it is not aimed at US public records. If you know of a vendor benchmark we have missed, send it and we will link it.
The tool these numbers came from
Everything above is the KeptPDF CLI running on its own machine. Install it with npm install -g keptpdf (Node 18.17 or newer, on Windows, macOS and Linux), and the first 100 files are a free trial, counted for the lifetime of that machine rather than per run. Redaction is keptpdf redact, with keptpdf review writing one row per candidate to a CSV for a person to approve first. Splitting is keptpdf split --auto, and --read-scans reads the scanned pages before it scores them. The document is opened by your own process, and the CLI makes no network call for it, so it runs on an air-gapped server.
A license is $4,800 a year for one organization, unlimited files, on your own hardware. See the CLI, or use the form on that page to talk to a person before you buy anything.
Questions, answered.
Why publish the numbers that make you look bad?
Will I get these numbers on my own documents?
Can I run these tests myself?
Why are finance and medical missing from the redaction table?
Does a splitting score in the fifties mean it cannot split my backfile?
--dry-run shows you the whole answer before anything is written.