Guide · Document separation

Split a scanned PDF into separate documents, with no separator sheets

Automatic document separation for a backfile that has already been scanned. The pages get read, not counted, and a person checks the proposed cuts before a single file is written.

You split a scanned PDF into separate documents automatically by reading the pages instead of counting them. The KeptPDF CLI command keptpdf split ./backfile -o ./documents --auto --read-scans reads a stack that has no text layer, works out where one document ends and the next begins, and writes one PDF per document named from that document’s own title. Nothing has to be inserted before scanning.

That is a different job from what the top results for this search actually do, and the difference is the whole reason this page exists.

Page one answers a question you did not ask

Search for this and you get splitters that cut a PDF every N pages, or at page ranges you type in yourself. That is a fine tool for a file where every document is the same length. A backfile is not that file. A property folder is a two-page permit, then a nineteen-page set of plans, then a one-page inspection card, then a deed. There is no N.

The rest ask you to name the ranges by hand, which is the exact work you were trying to avoid. Somebody still has to open the 800-page scan and find where page 47 stops being one thing and starts being another.

Separator sheets solve a problem you no longer have

The capture and scanning vendors do answer the right question. Most of the ones we checked answer it the same way: print barcode or patch-code sheets, or blank separator pages, and have someone insert them into the paper stack before it goes through the scanner. The scanner then cuts on the sheet.

That works on day-forward scanning, where you control the paper. It does nothing for a backfile that is already one big PDF. You cannot go back and insert a sheet into a scan that happened in 2019, or into a repository export, or into the file a records custodian just emailed you. Re-scanning the box to insert sheets is not a plan, it is a second project.

Separator sheets

The decision is made in the paper, before the scan. A backfile that is already a PDF has no place to put the sheet, so the answer arrives too late to use.

Read the pages

The decision is made from the file you already have. Nothing is inserted, nothing is re-scanned, and the same command works on a folder of 400 old scans.

Both approaches separate documents. Only one of them can be run on the box you already scanned.

What content-based separation does instead

--auto compares each page with the pages around it and proposes the points where the document changes. It then names each document from its own first page rather than from the stack it arrived in, so you get Property-Folder-01-building-permit.pdf instead of Property-Folder-page-1-to-10.pdf. Nothing gets an invented name: a document whose first page offers no usable title keeps its page range as a name.

The pages have to be readable for this to work well. A stack that came straight off a scanner has no text layer at all, which is why --read-scans exists: it reads every page that has no text before looking for the breaks, so the detector has words to compare instead of just ink. It costs minutes per file rather than seconds, so you ask for it rather than getting it by default, and it buys nothing on a PDF that already has text.

This is the same detector the KeptPDF website runs in its Organize panel, so a watch folder on a server and a clerk on the website cannot propose two different answers for the same packet.

Run it on your own backfile

  1. Install it. npm install -g keptpdf. It needs Node 18.17 or newer and runs on Windows, macOS and Linux, on Intel and ARM. There is no compiler and no build step.
  2. Look before you cut. --dry-run prints every document it found, its page range, and the reason each break was called, and writes nothing at all. Add --verbose and it also prints the breaks it nearly called, which is the difference between “raise the sensitivity” and “this really is one document”.
  3. Read the scans. Add --read-scans when the file has no text layer. A run that would have benefited from it says so at the end.
  4. Fix the list. --review holds everything back and walks you through the documents: rename one, join it to the one above, split it in two, skip it, or move the sensitivity dial and watch the list redraw. Nothing is written until you press w.
  5. Choose the shape. --output-shape files (the default) writes one PDF per document. bookmark writes one PDF with a bookmark per document. both writes both.
$ keptpdf split ./backfile -o ./documents --auto --read-scans --dry-run

+  Property-Folder.pdf  7 documents in 128 pages       4m 12s
     1. pages 1-2    Building Permit Application
     2. pages 3-21   Site Plan Set
     3. pages 22-24  Pages 22-24
     4. pages 25-25  Final Inspection Card

The detector proposes. It does not decide

Be clear-eyed about this, because a page that promises otherwise is lying to you. Automatic separation is a proposal. It will sometimes miss a seam, and it will sometimes call one where there is not one. Those two mistakes do not cost the same: an extra break costs you one keypress to merge, while a missed break has to be found first, which is the hard task you were paying to avoid. The default sensitivity leans toward finding breaks for exactly that reason, and --auto-sensitivity low | medium | high lets you move it. Medium is the default. Lower finds fewer and clearer breaks, higher finds more of them and some wrong ones.

That is what the --review step is for. It is not a formality bolted on for comfort, it is the step where a person who knows the records makes the call the software cannot. Measured accuracy for the detector, on named corpora with the misses shown rather than hidden, is on the benchmarks page.

When every page is the same city form

Here is the case that gets glossed over everywhere else. A Laserfiche property file is often a stack of one-page filings on the same city form: same header, same layout, same typeface, same margins, forty times in a row. Content-based separation reads how one page differs from the next, and on that stack almost nothing differs. It is the hardest shape in this whole category, and any tool that tells you it handles it cleanly has not tried it.

What actually helps on a form stack:

  • Turn the sensitivity up. On a stack where the pages barely differ you want more proposed breaks, not fewer, because merging is cheap.
  • Always use --review here. On a mixed packet review is a spot check. On a form stack it is the main event, and it is still far faster than opening the file and hunting.
  • Consider bookmarking instead of cutting. If the point is that somebody can find the right filing, a bookmark at each one does that without the risk of a wrong cut writing a wrong file.

One answer, three shapes

Finding the documents is one job. What you do with that answer is a separate choice, and cutting is not always the right one.

--output-shape bookmark writes <name>-bookmarked.pdf: every page of the input, in order, nothing removed and nothing reordered, with a top-level bookmark at the first page of each document carrying that document’s own title. No table of contents page is inserted, because that would push every page down by one and the page numbers already stamped on the packet would stop matching the file. A resident handed a 300-page scanned property folder wants to jump to the building permit, not to receive forty loose files. --output-shape both gives you the bookmarked file and the separate documents from the same detection.

--by-type adds the question a records office asks first: what kind of paper is this. Each document is filed into one of thirteen piles (permit, application, inspection, certificate, enforcement, decision, plans, survey, recorded, property card, submittal, correspondence, and other for anything it will not guess at). With separate files each document lands in a folder named for its pile, ready for a repository import. With the bookmarked shape the sidebar gets a heading per pile. In --review, t changes a pile before anything is written, and --json reports the pile and its confidence for every document.

Where the file goes, and what it costs

The PDF is read on the machine you run the command on. The CLI makes no network call for the document, and the OCR runs locally too, so an air-gapped workstation in a records office is a supported place to do this. That matters when the backfile is a property file, a personnel file, or a police record, and it is the same reason the redaction pass on a backfile job can run in the same room as the scanner.

Exit codes are scriptable: 0 done, 1 at least one file failed, 2 usage error, 3 license problem. A nightly job over a scanning bureau’s output can fail loudly instead of quietly producing nothing.

The trial runs on your first 100 files, as a lifetime total per machine rather than a per-run cap, with no account and no card. For split --auto, detection itself is never limited: the trial reads every page of a file of any length and lists every document it finds, because whether the boundaries are right is the thing you are evaluating. What it writes is capped at the documents that fit inside the first 50 pages. A --dry-run is never capped at all. A license is $4,800 a year for one organization, unlimited files, on your own hardware. Details and the buy card are on the CLI page.

Questions, answered.

Can it split a scan that has no text layer at all?
Yes, and that is the case to plan for. Add --read-scans so the pages with no text are read before the breaks are looked for, which gives the detector words to compare rather than ink alone. It costs minutes per file rather than seconds, and it buys nothing on a PDF that already has a text layer, so it is asked for rather than assumed.
Do I need to insert barcode or blank separator sheets first?
No. Separator sheets have to go into the paper before the scan, which is no help on a backfile that is already a PDF. This reads the file you already have, so nothing is printed, inserted or re-scanned.
How wrong can the proposed cuts be?
It can miss a seam and it can call one that is not there. The default sensitivity leans toward proposing breaks, because an extra break is one keypress to merge while a missed one has to be found by hand. Use --dry-run to look first and --review to fix the list before anything is written, and see the benchmarks page for measured accuracy on named corpora.
What about a stack of identical city forms?
That is the hardest shape in this category, and we will not pretend otherwise. When forty one-page filings share the same form, header and layout, there is very little difference between pages for content-based separation to read, so it finds fewer of the seams. Raise the sensitivity, always use --review on those files, and consider the bookmarked output shape so a wrong cut cannot write a wrong file.
Can I get one bookmarked file instead of dozens of loose PDFs?
Yes. --output-shape bookmark writes one PDF with every page in order, nothing removed and nothing reordered, and a top-level bookmark at the first page of each document. No table of contents page is inserted, so the page numbers already on the packet still match. --output-shape both writes the bookmarked file and the separate documents from the same detection.
Does the file get uploaded anywhere?
No. The PDF is read on the machine you run the command on, and the CLI makes no network call for the document. The OCR runs locally as well, so it works on an air-gapped workstation.
What does it cost, and can I try it on my own backfile?
The trial runs on your first 100 files, a lifetime total per machine rather than a per-run cap, with no account and no card. Detection is never limited: it reads a file of any length and lists every document it finds, and a --dry-run is uncapped. What the trial writes is capped at the documents inside the first 50 pages. A license is $4,800 a year for one organization, unlimited files, on your own hardware.

Point it at the backfile you already scanned.

See the KeptPDF CLI