Skip to main content
SYSTEM ONLINE · aws us-west-1 · build 26.04.2
POs & NET 30 accepted · 559-251-7767 · Fresno, CA
SentraCheck // document compliance
Login Request a Demo
Privacy • August 2026

Before You Publish the Archive: Finding PII in 40 Years of Documents

Scanning a backlog turns a filing cabinet into a website. Material that was practically obscure for thirty years becomes searchable by anyone, instantly.

There is a concept in privacy law usually called practical obscurity: a record can be technically public and still, as a practical matter, unavailable, because finding it requires knowing it exists and driving to a counter to ask for it.

Digitization ends practical obscurity. The same record, indexed and published, is one search away from anyone on earth, forever. That is largely the point — and it is exactly why the material inside those files deserves a look before the archive goes live.

Not legal advice. What must be withheld or redacted depends on the record, the applicable exemptions, and your counsel's judgment. This article is about finding the material so somebody qualified can decide.

What turns up in old files

Records created decades ago were created under different norms. Social Security numbers appeared on forms as routine identifiers. Home addresses and phone numbers were printed without a second thought. Nobody was thinking about a search engine.

The categories that most often surface in a historic municipal corpus:

  • Social Security numbers — on older permit applications, business licenses, claim forms, employment records, and benefit applications
  • Dates of birth and driver's license numbers — on applications and identification copies
  • Financial account and routing numbers — on vendor forms, direct deposit authorizations, and refund requests
  • Medical information — in claims, workers' compensation files, ADA accommodation requests, and ambulance or incident reports
  • Handwritten signatures — reproducible at high resolution from a scanned image
  • Home addresses of protected individuals — safety personnel and others whose addresses may be protected by statute
  • Information about minors — in program registrations, incident reports, and school-related correspondence
  • Personnel material — evaluations, discipline, and complaints, frequently exempt and frequently mixed into project files

Where it hides

Reviewing the obvious forms is not sufficient. Personal information migrates:

  • Attachments and exhibits. A clean staff report with an appendix of raw claim forms.
  • Agenda packets. Correspondence from members of the public, submitted with full contact details and reproduced into the packet.
  • Litigation and claim files. Discovery material swept into a project folder.
  • Email printed to PDF. Headers carrying distribution lists and personal addresses.
  • Document metadata. Author names, network paths that expose employee usernames, and in some formats an editing history.
  • Beneath a scanned image. Once OCR has run, the text is in the file even though the page looks like a picture. Searching only the visible layer misses it.

The OCR paradox. Adding a text layer is necessary for search and for accessibility — and it is what makes the Social Security number on page 412 of a scanned permit file findable by a search engine. Text extraction is exactly why you must scan for personal information before publishing, not after.

Black boxes are not redaction

The most common and most damaging mistake: opening a PDF, drawing a filled black rectangle over sensitive text, and saving. The rectangle is an annotation drawn on top. The characters underneath are untouched.

Anyone can recover them by selecting the text, copying it, or running a text extraction. This has produced real disclosures at agencies large enough to know better.

Real redaction removes the content — the text objects, the image regions, and any metadata that carries it — and then flattens the result. Use a tool with a genuine redaction function, apply it, and save a new file.

Then verify, every time:

  1. Extract the text of the redacted file (pdftotext redacted.pdf -) and search the output for the values you removed.
  2. Search the file for the requester's terms and for known patterns such as nine-digit sequences.
  3. Check the document properties for author names, paths, and comments.
  4. On scanned pages, confirm the image itself was altered, not just the text layer — if the number is still legible in the picture, nothing was redacted.

A workable process for a large archive

Reviewing 200,000 pages by hand is not realistic. Reviewing zero pages is not defensible. The middle path is automated detection followed by targeted human review.

  1. Extract text from everything, including scanned material, so the whole corpus is searchable.
  2. Run pattern detection across it — Social Security number formats, dates of birth, account and card numbers, driver's license patterns, and medical vocabulary. Expect false positives; a nine-digit parcel number looks a great deal like an SSN.
  3. Rank by density and by series. Personal information clusters. Whole record series will be clean; a handful will be saturated. That distribution tells you where to send human reviewers.
  4. Human review of the hits, plus a random sample of the apparent clean set, because pattern matching does not catch prose. "The claimant's daughter, age 7, attends Lincoln Elementary" matches no regular expression.
  5. Redact properly, then verify using the steps above.
  6. Record the decisions. What was withheld, under which basis, and who approved it. If the redaction is ever challenged, that record is your answer.

Decide what belongs online at all

Publication is a choice, not an obligation that follows automatically from digitization. Some series are worth scanning for internal use and retrieval while remaining available only on request. Older personnel files, claim files, and anything heavy with third-party personal information are common candidates for scan-but-do-not-publish.

That decision belongs to counsel and the records officer together, made deliberately and written down — not made by default because a folder was easy to upload.

Sequence it correctly: scan, OCR, detect personal information, review and redact, verify, then remediate for accessibility and publish. Detection has to come after OCR, because it needs the text — and it has to come before publication, because after publication the material has already been indexed by every crawler on the internet.

Scan the archive before the public does

SentraCheck detects Social Security numbers, dates of birth, financial account numbers, and other personal information across an entire document corpus, and shows you exactly where each one sits.

Request a Demo More Articles