Redact PDFs from the command line
One command over a directory, run on your own machine. Here is the install, the exact flags, what lands in the output folder, the approval step that runs before anything is burned, and how to put it on a schedule.
To redact PDFs from the command line, install the KeptPDF CLI from npm and point one command at a folder: keptpdf redact ./intake -o ./produced -r. It finds names, government IDs, account and case numbers, addresses, emails and phone numbers, burns them out of the page, re-reads every file it writes to confirm the text is gone, and drops a certificate beside each one. Nothing is uploaded.
This page is the how-to. The product page, with pricing and the buy form, is /cli.
Why the usual answers stop short
Search this and you get two kinds of result. The first is a handful of small open-source scripts. The ones we looked at take a list of words you already know and remove those words. That is useful when you know the words. It is not a redaction pass over a backfile, where the whole problem is that nobody knows which names are sitting in file 3,880.
The second is a desktop editor walkthrough. Those are per-file by design: open a document, run a search, apply the redactions, save, repeat. A records office with a folder of 4,000 releases does not have 4,000 openings in it.
Both leave the same two jobs undone. Nothing checks the finished file to prove the text is actually gone, and nothing leaves a record of what was removed. Those are the two things a production gets questioned on later.
Install
$ npm install -g keptpdf $ keptpdf --version
Or skip the install and run it once with npx keptpdf redact ./intake -o ./produced -r.
It needs Node 18.17 or newer and runs on Windows, macOS and Linux, on Intel and ARM. That is the whole requirement list. It installs with two dependencies and 27 packages, no compiler and no build step, and the PDF and OCR engines ship inside the package, so what you install is what was tested.
The command, over a directory
# every PDF in ./intake, including subfolders $ keptpdf redact ./intake -o ./produced -r ✓ 2024-intake-0417.pdf 14 boxes verified ✓ 2024-intake-0418.pdf 9 boxes verified ✓ scanned-packet-22.pdf 31 boxes verified (OCR'd first) ✓ 2024-intake-0419.pdf 11 boxes verified 418 redacted, 0 skipped, 0 failed
Inputs can be files, directories or simple globs such as scans/*.pdf. A directory takes its top-level PDFs unless you add -r, and with -r the output mirrors the input tree instead of flattening it. A page with no text layer is read with OCR on the same machine first, then redacted like any other page. Word, Excel, PowerPoint and OpenDocument files are accepted too: they are converted to PDF locally first, which needs LibreOffice installed.
| Flag | What it does |
|---|---|
-o, --out <dir> | Output directory. Required unless you are using --dry-run. |
-r, --recursive | Descend into subdirectories. The output mirrors the tree. |
--dry-run | Detect only. Prints counts, writes nothing. |
--sensitivity <level> | balanced (default), conservative or thorough. |
--only <cats> / --skip <cats> | Keep or drop categories by id. keptpdf categories lists them. |
--terms <file> / --ignore <file> | Words to always redact, and detected values to never redact. One per line. |
--label <code> / --labels <file> | Stamp an exemption code inside each box, one for everything or one per category. |
--log <file> | Write a redaction log of every box to .csv or .json. |
--apply <file.csv> | Burn the decisions from a reviewed file instead of the sensitivity default. |
--searchable | Read a text layer back off the finished redacted pages so the output stays searchable. |
--suffix <str> / --overwrite | Change the output name suffix (default -redacted), or redact over an existing output. |
--json | Print one JSON document to stdout instead of the lines above. |
What lands in the output folder
For each input, three or four files come out:
- The redacted PDF, named
<name>-redacted.pdf. Set a different suffix with--suffix, or pass an empty one to keep the original name. .cert.txtand.cert.jsonbeside it: the counts by category, the SHA-256 of the input and of the output, and the verification result. A branded.cert.pdfis written on licensed runs.--no-certturns the certificates off.- The redaction log, if you asked for one with
--log production.csv. One row per box: file, page, category, code, method, confidence, position and the certificate id. It carries no redacted text on purpose, because the log travels with the production and listing what you hid would undo the redaction.
The check that matters happens between writing and finishing. Every output is re-opened and every string in it is compared against what was supposed to be gone. If anything survived, the file is not written at all and the run reports a failure. That is the part a black rectangle drawn over live text can never do, and it is the mistake that keeps making the news.
Just the SSNs, across thousands of files
State statutes that require Social Security numbers to be masked before release do not care how many files there are. Narrow the run to one category and look before you burn:
$ keptpdf redact ./land-records --only ssns --dry-run -r $ keptpdf redact ./land-records -o ./produced --only ssns -r --log ssn-pass.csv
keptpdf categories prints every id you can pass to --only and --skip, including names, ssns, addresses, accountNumber and dates. A dry run writes nothing, so it is safe to point at a whole repository export while you work out what the pass should cover.
A person approves before anything burns
Redaction is not reversible, and automatic detection is a first pass rather than a promise of completeness. So the CLI can hand you a spreadsheet first.
$ keptpdf review ./intake -o decisions.csv
# open decisions.csv, change any wrong "keep" to "redact" (or the reverse)
$ keptpdf redact ./intake -o ./produced --apply decisions.csv
One row per candidate, with the page, the category, the confidence and the detected text. Rows above the sensitivity threshold arrive marked redact, rows below it arrive keep. An untouched file applies exactly what a plain run would have done, so adding the review step cannot quietly stop redacting something. Each row also has a code column, so one document can cite one exemption on one bar and a different one on the next.
The review file holds the detected text, because a person has to read it to judge it. Treat it as a working file: keep it with the sources, and never ship it with the production.
On a schedule, or inside a script
Exit codes are the contract: 0 done, 1 at least one file failed, 2 usage error, 3 license problem. A nightly job can fail loudly instead of quietly producing nothing.
#!/bin/sh keptpdf redact /srv/intake -o /srv/produced -r --log /srv/logs/$(date +%F).csv if [ $? -ne 0 ]; then mail -s "redaction pass had failures" records@example.gov < /dev/null fi
Add --json and the run prints one JSON document to stdout instead of the progress lines: every file with its input and output path, status, page count, redaction counts by category, and the verification result. That is what you parse in a pipeline.
For a folder that fills up all day, use keptpdf watch <in-dir> -o <out-dir>, which redacts each new file as it arrives. Add --once to process what is there now and exit, which is the shape a cron job or a scheduled task wants. Saving a watch from the interactive menu goes one step further: it writes a ready-to-install service file for the machine it is on, a systemd unit on Linux, a Task Scheduler command on Windows, or a launchd plist on macOS, so the folder is served from boot with nobody logged in.
How this differs from calling a cloud redaction API
The difference is not the redaction. It is where the document goes to get redacted. A hosted API has to receive the file first, which adds a party to the chain of custody for exactly the documents where the chain of custody is the point: personnel files, medical records, police reports, benefits claims, discovery. That is a procurement conversation, a data-processing agreement and a retention question before a single page is redacted.
The KeptPDF CLI is opened, run and written by a process on your own machine. It makes no network call for the document, and none for its license either, which is checked with an offline signature, so a batch that runs overnight cannot be stopped by someone else's server being down. It runs on an air-gapped box. The one exception is opt-in and yours to invoke: keptpdf update asks npm's public registry for the newest version number and then installs it, and only when you run it.
Pricing shape differs too. A per-page or per-document API meters volume; a license does not, so a 50,000-page backfile costs the same as a 500-page one. There is a fuller side-by-side on the CLI page, and measured accuracy numbers, with the misses shown, on /benchmarks.
One honest limit, since people searching for a self-hosted redaction API are usually picturing an HTTP endpoint. This is a command, not a server. If you need an endpoint, you wrap the command in your own service, or you point a watched folder at the queue your application already writes to. We do not ship a daemon that listens on a port.
Try it, then license it
The trial covers your first 100 files, with no account and no card. That is 100 files in total on that machine, not per run, and there is no deadline to spend them. The count lives in a plain file on your own machine. Trial output carries one small footer line naming it as trial output, stamped into the page as an image, so it is fine for judging the redaction on your own documents and not fit to file or produce. A license key removes it.
A license is $4,800 a year for one organization, unlimited files, running on your own hardware. Never per page, never per seat. The buy card is at /cli#buy, and there is a form on the same page if you would rather talk to a person first about volume, a pilot or procurement steps.