Backfile conversion: document separation and a redaction pass
Two steps you can add to a conversion job at software cost instead of labor cost. Separate the documents inside a stack that has already been scanned, then redact the batch before it goes into the client’s repository.
Backfile conversion software that does document separation reads the scanned page stream and works out where one document ends and the next one begins, with no separator sheets and no barcodes inserted before scanning. The KeptPDF CLI does that on a stack you have already scanned, and adds a bulk redaction pass before repository import. Both run on your own machines.
That matters because the two hardest parts of a backfile job are the two you cannot bill as scanning: finding the document boundaries in a page stream, and getting personal data out before the client’s records go live in a public-access portal.
Where the two steps sit in a job you already sell
A conversion job has a shape: pick up the boxes, prep, scan, OCR, index, QC, import into the client’s ECM, hand the paper back or destroy it. Separation and redaction are the two places where the work quietly stops being machine time and starts being a person at a screen.
- Separation gets decided before the scanner runs. Someone prints separator sheets or patch code pages and inserts them by hand between documents. That is fine on a day-forward job. On a backfile that is already scanned, or on files a client sends you as PDFs, the decision has been missed and cannot be made again without rescanning.
- Redaction gets quoted as a manual line item, priced per page or per hour, because the alternative is sending client records to a cloud service. For land records, court files, personnel files and permit folders, that upload is often the thing the client is trying to avoid.
Both steps become a command you run on a folder. The billable step stays on the invoice; the labor behind it does not.
Separation on a stack that is already scanned
keptpdf split with --auto reads every page, works out where each document starts, and writes one PDF per document named from that document’s own title rather than from the stack it arrived in.
# look first, write nothing $ keptpdf split ./backfile -o ./documents --auto --read-scans --dry-run + Property-Folder.pdf 7 documents in 214 pages 1. pages 1-18 Building Permit Application 2. pages 19-33 Certificate of Survey 3. pages 34-96 Inspection Records
On an image-only backfile, add --read-scans. It reads every page that has no text layer before it looks for the breaks, so the detector has the words instead of only the ink. It costs minutes per file rather than seconds, which is why it is asked for rather than assumed, and it buys nothing on a PDF that already has text.
Without that flag, scanned pages are still not skipped. A page with no text layer is fingerprinted by the shape of its ink: how dark the paper is, where the margins fall, how far apart the lines sit. Only the pages that need it are rendered, so a born-digital file pays nothing for it. --auto-sensitivity low | medium | high decides how hard it looks, and medium is the default: lower finds fewer, clearer breaks, higher finds more of them and some wrong ones.
--by-type answers the question a records office asks first: what kind of paper is this. Each document is dropped into one of thirteen piles (permit, application, inspection, certificate, enforcement, decision, plans, survey, recorded, property-card, submittal, correspondence, and other for anything it will not guess at) and written into a folder named for its pile. That is an index column you can hand the client, produced by the same pass.
What the output looks like
One detection, two shapes, chosen with --output-shape.
filesis the default: one PDF per document, named from its own title, ready to import as separate records.bookmarkcuts nothing. It writes<name>-bookmarked.pdf: every page of the input, in order, nothing removed and nothing reordered, with a top-level bookmark at the first page of each document. No table of contents page is inserted, so the page numbers already printed on the packet still match the file.bothwrites the separate files and the bookmarked copy.
The bookmarked shape is the one clients ask for when the record should stay whole. A resident handed a 300-page scanned property folder wants to jump to the building permit, not to receive forty loose files.
The redaction pass before import
Once the stack is separated, the redaction pass runs over the folder.
$ keptpdf redact ./documents -o ./produced -r --label "(b)(6)" --log production.csv ✓ property-folder-01-building-permit.pdf 14 boxes verified ✓ property-folder-02-certificate-of-survey.pdf 9 boxes verified 214 redacted, 0 skipped, 0 failed
Four things about that run matter to a bureau, because they are the four things a client will ask about later.
- It burns rather than covers. The characters are removed from the content stream and the area is painted. A black rectangle drawn on top of live text is not a redaction.
- It re-reads its own output. After writing, the file is re-opened and every string in it is checked against what was supposed to be gone. If anything survived, no file is written and the run reports a failure. That is what turns a batch into something you are willing to sign off on.
- Every file leaves evidence. Each output gets a
.cert.txtand a.cert.json, and a branded.cert.pdfon licensed runs: the counts by category, the SHA-256 of the input and of the output, and the verification result. - The log travels with the production.
--logwrites one row per box: file, page, category, code, confidence, position and the certificate id. It carries no redacted text on purpose, because a log that lists what you hid would undo the redaction.
Scanned pages are handled in the same pass: a page with no text layer is read on the machine first, then redacted like any other. If the client needs the delivered files searchable, --searchable adds a text layer back by reading the finished redacted pages, after the verification pass, so the layer cannot contain what was removed.
--label stamps an exemption code inside each box, clipped to it, never growing the redaction. --labels <file> gives a different code per category. KeptPDF ships no exemption preset, because which code applies is the client’s legal call, not a vendor default.
The operator review step
Redaction is not reversible, and automatic detection is a first pass, not a promise of completeness. So both halves of the job have a human gate you can staff with one reviewer instead of a room of them.
$ keptpdf review ./documents -o decisions.csv
# open decisions.csv, change any wrong "keep" to "redact" (or the reverse)
$ keptpdf redact ./documents -o ./produced --apply decisions.csv
The review file is one row per candidate, with the page, the category, the confidence and the detected text. Rows above the sensitivity threshold arrive marked redact; rows below it arrive keep. An untouched file applies exactly what a plain run would have done, so adding the review step can never quietly stop redacting something. It holds the detected text, because a person has to read it to judge it: keep it with the sources and never ship it with the production.
Splitting has its own gate. --review holds everything back and walks the list: rename a document, join it to the one above, split it in two, change its pile, or leave it out. Nothing is written until you say write.
Exit codes are scriptable, so this drops into the job scheduler you already run: 0 done, 1 at least one file failed, 2 usage error, 3 license problem. A nightly batch fails loudly instead of quietly producing nothing.
It runs on your hardware, which is half the sale
The CLI makes no network call to do the work, and none to check its license, which is verified with an offline signature. It runs on an air-gapped box. The one exception is opt-in and yours to invoke: keptpdf update asks npm for the newest version number and installs it, only when you run it.
For a bureau, that removes a conversation rather than documenting one. A cloud redaction service adds a party to the chain of custody for the exact records where the chain of custody is the point, and a county attorney will ask about it.
What it costs on a job
The license is $4,800 a year, one organization, unlimited files, your own hardware. Never per page and never per seat. You license it once and run it on every client job, which is what makes the arithmetic work: on a conversion contract the redaction line has been priced as labor, and the software cost does not move when the page count does.
Try it first. The trial covers your first 100 files on a machine, as a lifetime total rather than a per-run cap, with no account and no card. Trial output carries one small footer line, so you can judge the redaction on real client files. Split detection is never capped: the trial reads a file of any length and lists every document it finds, and saves the ones inside the first 50 pages.
$ npm install -g keptpdf $ keptpdf split ./backfile -o ./documents --auto --read-scans --dry-run
It needs Node 18.17 or newer and runs on Windows, macOS and Linux, on Intel and ARM. That is the requirement list.
If you resell, or want a pilot first
Solution providers can buy the license on a client’s behalf. Each client gets its own key, signed to their name, and every certificate the tool produces names that organization. The partner code takes 25% off the first year ($3,600 instead of $4,800); you bill your client at your own rate and keep the difference, and renewals are at list price.
The /cli page has two doors for that reason. The buy card takes an organization name and an email straight to payment, and the key is emailed when payment clears. The second card is a short form that goes to a person instead: your organization, your email, what you need (a quote or a PO, a pilot on your own files, the partner code, or something else) and an optional note about volume, deadline or procurement steps. It sends only what you type, never a document, and a human replies from support@keptpdf.com.