Guide · Public records

How to redact PII for a public records request

What comes out is your statute’s call. Getting it out for good, putting the exemption code on each box, and keeping a log that travels with the production is the part a tool can carry.

To redact PII for a public records request, decide which exemption covers each item, remove that text from the file rather than covering it, stamp each box with the exemption you are citing, and keep a log of every box you drew. KeptPDF’s command line does those last three across a whole folder and checks each finished file.

This page is about method, not law. Nothing here is legal advice, and which exemptions apply to your records comes from your own statute and your own counsel.

What usually has to come out

Across most records offices the same items keep showing up on the redaction list:

  • Social Security numbers, and the state ID numbers that sit beside them: driver’s license, passport, tax IDs.
  • Dates of birth, which are usually the second thing an identity thief needs after a name.
  • Bank account and routing numbers, card numbers, and the account or member IDs on a utility bill or a benefits file.
  • Home addresses and personal phone numbers, often only for certain classes of person, which is where the statutes differ most.
  • Medical information in a personnel file, a first responder report, or an accommodation request.
  • Third-party personal information: the neighbour who complained, the witness, the child on the incident report. More on that below.

That list is a starting point, not a rule. The exact set depends on your state’s public records statute and on the specific exemption you are relying on for that record. Some states exempt a home address only for named categories of employee. Some require you to release the record with a name struck out rather than withhold it. We will not cite statute numbers here: the one that matters is the one your office already works under.

So the tool should not pick for you. In the KeptPDF CLI, keptpdf categories prints the category ids this build looks for, and --only and --skip narrow a run to the ones your policy covers. --terms <file> takes a list of phrases, one per line, that always come out, which is how a specific person’s name gets swept consistently through an entire release. --ignore <file> is the reverse: exact detected values that should never be struck, so the mayor’s published office address stops being flagged on every page.

A black box over text is not a redaction

This is the mistake that keeps making the news. A rectangle drawn in a PDF editor is a shape sitting on top of a page. The characters are still in the file underneath it. Select all, copy, paste, and the text you hid comes back out. So does a text extraction tool, and so does the requester who knows the trick.

A drawn box

The rectangle is a shape on top of the page. The words are still in the file. A copy and paste, or any extraction tool, reads them straight back out.

A burned redaction

The characters are taken out of the page content and the area is painted. The finished file is re-opened and re-read before it is handed over.

Both look identical on screen. Only one of them survives a requester with a mouse.

KeptPDF removes the characters from the page content and paints the area, then it does the thing a records office actually needs: it re-opens the file it just wrote and re-reads it. If any of the text that was supposed to be gone can still be pulled out, no file is written at all and the run reports a failure. You cannot accidentally hand out a bad file, because a bad file does not get produced.

Two details matter for records work. A page with no text layer, which is most of a scanned property file, is read with OCR first and then redacted like any other page. And redacted pages are flattened to images, which is safe but not searchable; --searchable puts a text layer back by reading the finished redacted pages, after that check, so the new layer cannot carry what was removed.

The exemption code belongs on the box

Most production protocols want each black bar to say why it is there. In the CLI that is one flag:

$ keptpdf redact ./records -o ./produced --label "(b)(6)" --log production.csv

--label stamps the code inside the box, clipped to it, so citing an exemption never makes the redaction bigger than what you redacted. --labels <file> gives a different code per category, which is how a run can cite one exemption on the SSNs and another on the third-party names. And a code column in the review file sets it box by box, so a single document can cite one exemption on one bar and a different one on the next.

KeptPDF ships no exemption preset. The interactive menu offers the federal 5 U.S.C. 552(b) list as suggestions with nothing preselected, and any code you type by hand works exactly the same, including a state statute cite or a paragraph of a protective order. Which one applies is your call, not our default.

The log that travels with the production

--log production.csv writes one row for every box in the run. The columns are fixed: file, page, category, code, method, confidence, x, y, width, height, and the certificate id of the file the box lives in. Position is stored as a fraction of the page, so a row locates a bar without anyone needing your original.

The log deliberately does not record the redacted text. It is meant to travel with the production, and a log that lists what you hid would undo the redaction it documents. Location, category and code is what the reader of a log needs. Name the file .json instead of .csv and it writes JSON with the same fields, the easier shape if your request tracking system will ingest it.

Be clear about what this is. It is the mechanical index of every redaction you made, not a Vaughn index or a narrative justification. It gives you the coordinates and counts behind whatever narrative you write, and an answer when someone asks six months later which page of which file carried which claim.

Records that mention people other than the requester

The hardest part of most requests is not the SSN. It is the third party: the complainant named in an inspection file, the witness in a police narrative, the other parent in a school record. The record is responsive and you have to release it, but the person mentioned in it never asked to be in a public production.

That is a judgment call on every hit, so the CLI has a step that hands you the calls before anything is permanent:

$ keptpdf review ./records -o decisions.csv
# open decisions.csv, flip any wrong "keep" to "redact" (or the reverse)
$ keptpdf redact ./records -o ./produced --apply decisions.csv --log production.csv

The review file is one row per candidate with the decision, the file, the page, the category, the confidence, a code column for the exemption, and the detected text. Rows above the sensitivity threshold arrive marked redact; rows below it arrive keep. If you hand the file back untouched, the run applies exactly what a plain run would have done, so adopting the review step can never quietly stop redacting something.

Two warnings belong in your desk procedure. The review file holds the detected text unredacted, because a person has to read it to judge it: keep it with the sources and never ship it with the production. And automated detection is a first pass. It can miss things, especially handwriting and smudged scans, so a person still has to look.

A certificate per file, for the day someone asks

Every output gets a .cert.txt and a .cert.json beside it, and a branded .cert.pdf on licensed runs. Each one records the counts by category, the SHA-256 of the file that went in and of the file that came out, and the result of that verification pass. On a licensed run the certificate also names the organization the license is signed to, on the document and on its verify page, so a deliverable shows whose license produced it.

That is a record of what your office did, and proof the produced file has not changed since. It is not a legal opinion or a promise about how anyone will rule on your exemptions.

One request, or a whole backlog

For a single file, the browser redact tool is free and does the work on your own device, so the document is not uploaded anywhere to be processed. That is the right tool for the request that arrived this morning with one PDF attached.

For a folder, a monthly release, or a backfile someone has been putting off, the command line is the one that scales. It takes files, directories or globs, recurses with -r, and keptpdf watch turns a folder into a drop box that redacts each new file as it lands, which is how a small office gives its intake clerk a one-step process. Exit codes are scriptable (0 done, 1 at least one file failed, 2 usage error, 3 license problem), so a scheduled overnight job fails loudly instead of quietly producing nothing.

The whole thing runs on your machine. The document and its text never leave it, and the CLI makes no network call for the work or for the license check, so it runs on an air-gapped box in a records room.

Try it on your own request

$ npm install -g keptpdf
$ keptpdf review ./records -o decisions.csv
$ keptpdf redact ./records -o ./produced -r --apply decisions.csv \
    --label "(b)(6)" --log production.csv

The trial runs on your first 100 files on that machine, with no account and no card, and trial output carries one small footer line so you can judge the redaction on real records first. A license is $4,800 a year for one organization, unlimited files on your own hardware, with no per-seat charge and no seat count to track. It needs Node 18.17 or newer and runs on Windows, macOS and Linux, on Intel and ARM.

See the full command reference, the buy card and the form for a quote or a pilot on the KeptPDF CLI page, and the measured accuracy numbers with the corpora they came from on benchmarks.

Questions, answered.

Does KeptPDF decide which exemption applies to my records?
No, and it ships no exemption preset. The interactive menu offers the federal 5 U.S.C. 552(b) list as suggestions with nothing preselected, and any code you type by hand works the same, including a state statute cite or a protective order paragraph. Which exemption covers a given item comes from your statute and your counsel, not from a default in our software.
Can a requester copy the text back out from under the black boxes?
Not from a file KeptPDF wrote. The characters are removed from the page content rather than covered, and the finished file is then re-opened and re-read. If any of the text that was supposed to be gone can still be extracted, no file is written and the run reports a failure. What a person still has to check is whether everything that should have been marked was marked.
What is in the redaction log, and could it leak what I hid?
One row per box, with the file, the page, the category, the exemption code, the method, the confidence, the position on the page, and the certificate id. It carries no redacted text on purpose, because the log is meant to travel with the production and listing what you hid would undo the redaction. The separate review file does hold the detected text, so keep that one with your sources and never ship it.
Our older records are scanned images with no text layer. Does this work on them?
Yes. A page with no text layer is read with OCR on your machine first, then redacted like any other page. Redacted pages come out flattened to images, and the --searchable flag adds a text layer back by reading the finished redacted pages after the verification pass, so the new layer cannot contain what was removed. OCR is not perfect on a poor scan, which is another reason a person reviews.
What does it cost for a records office?
$4,800 a year for one organization, unlimited files on your own hardware, with no per-seat charge and no seat count to track. The trial runs on your first 100 files on a machine with no account and no card, and trial output carries one small footer line. Redacting a single PDF in the browser at /redact-pdf is free.

Run one request through it, then a folder.

See the KeptPDF CLI