Guide · Records offices

Laserfiche redaction workflow: bulk redaction, before or after import

Quick Fields and Workflow can redact what matches a pattern you configured. Here is what that leaves a person to do by hand, and the pull, redact, push loop for the documents already sitting in the repository.

Laserfiche redacts inside its own tools: Quick Fields Auto-Annotation redacts pattern matches at capture, and Workflow’s Apply Text Annotation activity redacts text found by a regular expression. For the bulk work those two do not cover, the practical loop is to pull the PDFs out through the repository API, run one redaction command on your own machine, and import the results back.

Where redaction sits in a Laserfiche intake

There are two jobs here, and they need different machinery.

Before import. Day-forward scanning. Paper comes in, gets captured, and lands in the repository. If the redaction happens in this lane, the sensitive text never gets filed in the first place.

Already in the repository. The backfile: twenty years of property files, personnel folders and council packets imported long before anyone wrote a redaction rule. This lane produces the release copies for public records requests, and it is where the manual work piles up.

A rule you configure at capture time does nothing for the second lane. That is the gap most of the unanswered Laserfiche Answers threads are actually about.

What Laserfiche’s own tools do

Quick Fields handles capture-time processing. Its Auto-Annotation process can be configured to look for a pattern of digits and dashes that fits a Social Security number and redact every instance that matches that pattern. Its Fixed Annotation process puts an annotation, redaction included, at a specified location on each document, which suits a form where the box is always in the same place.

Workflow handles documents in the repository. The Apply Text Annotation activity hides, partially hides or emphasizes text on a document, using regular expressions to find it. It can apply a redaction, a highlight, a strikethrough or an underline, and Laserfiche’s documentation notes that creating a redaction needs Laserfiche Server 10.3 or later.

The client handles the rest by hand. A person opens the document in the viewer, picks the redaction tool and drags a box, and may be prompted for a reason from a configured list. Laserfiche describes the result as an annotation: users who hold the See Through Redactions entry access right can view the material underneath, and can choose whether to keep the redactions when they export or print.

Laserfiche Cloud is the exception to check for yourself. When we worked through the Cloud developer documentation in August 2026 we found no built-in auto-redaction step there. What Cloud does have is a clean REST API for getting documents out and putting them back, which is what makes the loop below possible.

What Workflow and Quick Fields leave to a person

Four things.

  • Anything without a shape. A regular expression is a shape. A Social Security number has one. A resident’s name does not, and neither does a street address written into a sentence, or a date of birth in a paragraph of narrative. Those get read by a person.
  • Pages with no text. A Laserfiche Answers thread on the prerequisites for automated redaction puts it directly: the documents would need to be stored as native Laserfiche documents (TIFF pages) and fully text-searchable, and OCR’d first if they lack text on every page.
  • The review call. The same thread recommends a human review step afterward, because OCR is not perfectly accurate and automated redaction can miss something handwritten. That advice is right, and true of our tool too.
  • The copy that leaves. Inside the repository, an annotation plus an access right is a sensible control. A release copy is a different artifact: a file that goes to a requester, with no rights model following it. What you want in that file is characters that are not there any more.

The loop: pull, redact, push

Three steps. Laserfiche keeps custody at both ends; the redaction happens on a machine you control.

One-time setup takes about ten minutes and needs admin rights on the Laserfiche account:

  1. Service principal. In Account Administration, create a service principal user with read rights on the input folder and write rights on the output folder, plus a service principal key.
  2. Service app. In the Developer Console, create a new app of type Service and point it at that service principal. Note the Client ID.
  3. Scopes. On the Authentication tab, enable repository.Read repository.Write. They are case-sensitive and space-delimited.
  4. Credential. On the same tab, create a long-lasting Authorization Key. That single string is the static secret your script uses. It dies if the service principal key is rotated, so plan rotation.

Then three API calls do the pulling. These are the US region hosts; swap in the Canadian or European hosts for those tenants, and check them against your own tenant before you rely on them.

# 1. a token (12 hour lifetime)
POST https://signin.laserfiche.com/oauth/token
  Authorization: Bearer <AUTHORIZATION_KEY>
  grant_type=client_credentials&scope=repository.Read repository.Write

# 2. list the input folder
GET  /repository/v2/Repositories/{repoId}/Entries/{folderId}/Folder/Children

# 3. export one document, then fetch the link it returns
POST /repository/v2/Repositories/{repoId}/Entries/{entryId}/Export
     {"part": "Edoc"}

The synchronous export times out at 60 seconds, so use ExportAsync with task polling for very large files. Write each PDF into a local folder. Now the middle step, which is one command:

$ npm install -g keptpdf
$ keptpdf redact ./in -o ./out -r

  ✓ 41827-parcel-file.pdf      22 boxes   verified
  ✓ 41828-parcel-file.pdf      17 boxes   verified
  ✓ permit-packet-2004.pdf     31 boxes   verified (OCR'd first)

  318 redacted, 0 skipped, 0 failed

A page with no text layer is OCR’d first, on the same machine, then redacted like any other page. The characters are removed from the content stream and the area is painted, rather than a rectangle being drawn over live text. After writing, the output is re-opened and every string in it is checked against what was supposed to be gone. If anything survived, no file is written and the run reports a failure. Each output gets a .cert.txt and a .cert.json, plus a branded .cert.pdf on licensed runs: the counts by category, the SHA-256 of the input and of the output, and that verification result.

The document does not leave the machine the CLI runs on. This recipe talks to exactly one service: your own Laserfiche tenant.

Exemption codes and a log that travels with the production

A public records production usually wants each black bar to say why it is there.

$ keptpdf redact ./records -o ./produced --label "(b)(6)" --log production.csv

--label stamps the code inside the box, clipped to it, so the redaction never grows to fit the text. --labels <file> gives a different code per category. The log records one row per box: file, page, category, code, method, confidence, position and the certificate id. It carries no redacted text, on purpose: the log travels with the production, and listing what you hid would undo the work. KeptPDF ships no exemption preset. Which code applies is your legal call.

The review step, which you should keep

Automatic detection is a first pass. It can miss things, and we will not claim otherwise, so the tool can hand you a spreadsheet before anything burns:

$ keptpdf review ./in -o decisions.csv
# open decisions.csv, change any wrong "keep" to "redact" (or the reverse)
$ keptpdf redact ./in -o ./out --apply decisions.csv

One row per candidate, with the page, the category, the confidence and the detected text. Rows above the sensitivity threshold arrive marked redact, rows below it arrive keep. An untouched file applies exactly what a plain run would have done, so adopting the review step cannot quietly stop something being redacted. A code column sets the exemption code per box, so one document can cite (b)(6) on one bar and (b)(7)(C) on the next.

The review file holds the detected text, because a person has to read it to judge it. Keep it with the sources and never ship it with the production. What we do claim, and prove per file, is that whatever you decided to redact is genuinely gone from the output.

Pushing the results back

Import is a multipart POST to the v1 endpoint, one call per file:

POST /repository/v1/Repositories/{repoId}/Entries/{outFolderId}/{name}?autoRename=true
     part "electronicDocument": the PDF bytes
     part "request": {"metadata": { ...optional template fields... }}

The certificates import alongside the documents, so the repository holds the proof next to the deliverable, and normal Laserfiche security, workflow and records management take over. On a self-hosted deployment the same shape works against the API Server’s own base URL and authentication.

The day-forward half: a watched folder

For scanning still arriving, you do not need the pull at all. Point the scanner output at a watched folder and import from the other side:

$ keptpdf watch ./scan-drop -o ./ready-to-import

Every PDF dropped into that folder comes out redacted, with its certificate. The interactive menu (type keptpdf with no arguments) can save a watch setup and write a ready-to-install service file: a systemd unit on Linux, a Task Scheduler command on Windows, a launchd plist on macOS. It then runs from boot with nobody logged in, and the import step watches the output folder.

A backfile that arrives as one long scan needs splitting first: see that guide.

What it costs, and how to check it before you commit

Install with npm install -g keptpdf. It needs Node 18.17 or newer and runs on Windows, macOS and Linux, on Intel and ARM. The trial covers your first 100 files with no account and no card. That is 100 files in total on that machine, not per run, and there is no deadline. Trial output carries one small footer line naming it as trial output, so it is not fit to file or produce, but it is enough to judge the redaction on your own documents.

A license is $4,800 a year for one organization, unlimited files, on your own hardware. For a solution provider, the license names the client organization and every certificate prints that name, so one client’s key does not fit another client’s work. Partner terms are on the CLI page, which also has a form if you would rather talk to a person first.

Measured accuracy, the corpora it was measured on, and the misses are on the benchmarks page.

KeptPDF is not affiliated with, endorsed by or sponsored by Laserfiche. Laserfiche is a trademark of Laserfiche. The descriptions of Quick Fields, Workflow and the Laserfiche client above come from Laserfiche’s own published documentation and a public Laserfiche Answers thread. Check them against the version you run, because product behaviour changes between releases.

Questions, answered.

Can Laserfiche redact documents automatically?
In places, yes. Per Laserfiche’s documentation, Quick Fields Auto-Annotation can be configured to find a pattern of digits and dashes that fits a Social Security number and redact every match at capture time, and Workflow’s Apply Text Annotation activity applies a redaction to text found by a regular expression, which needs Laserfiche Server 10.3 or later. Both work from patterns you configure, so what they leave behind is the material that has no fixed shape, such as a resident’s name or an address written into a sentence.
What tools are required to auto redact in Laserfiche?
Quick Fields for the capture lane and Workflow for documents already in the repository, plus the pages themselves being readable. A Laserfiche Answers thread on this question states that the documents would need to be stored as native Laserfiche documents (TIFF pages) and fully text-searchable, and OCR’d first if they lack text on every page. The same thread recommends a human review step afterward, because OCR is not perfectly accurate and an automated redaction can miss something handwritten.
How do I bulk redact documents that are already in the repository?
Export them through the repository API into a local folder, run one redaction command over that folder, and import the results into an output folder. In Laserfiche Cloud that is a service principal and a Service app with the repository.Read and repository.Write scopes, then a token call, a folder listing, an export per document, and a multipart import per result. The redaction step is keptpdf redact ./in -o ./out -r. The documents stay inside your own network the whole time.
Can I redact before importing into Laserfiche instead?
Yes, and for day-forward scanning it is the simpler shape. Run keptpdf watch on the folder your scanner writes to, with its output folder as the thing your import picks up. Every PDF dropped in comes out redacted with its certificate beside it, and the certificates import alongside the documents so the repository holds the proof next to the deliverable. The interactive menu can install the watch as a boot-time service on Windows, macOS or Linux.
Is KeptPDF a Laserfiche product or add-on?
No. KeptPDF is not affiliated with, endorsed by or sponsored by Laserfiche, and Laserfiche is a trademark of Laserfiche. KeptPDF is a command line tool that reads and writes PDF files on your own machine. It pairs with a Laserfiche repository through Laserfiche’s own REST API, the same way any script that exports and imports documents would.

Try it on 100 of your own files before you buy anything.

See the KeptPDF CLI