Skip to main content
SYSTEM ONLINE · aws us-west-1 · build 26.04.2
POs & NET 30 accepted · 559-251-7767 · Fresno, CA
SentraCheck // document compliance
Login Request a Demo
Records • August 2026

How to Inventory Every Document Your Agency Owns (Before You Archive Anything)

You cannot decide what to archive, what to scan, or what to remediate until you know what you have. Most agencies have never had a single list. Here is how to build one.

Every records project an agency runs — a retention cleanup, a scanning backlog, an ADA remediation program, a website migration — begins with the same question, and almost nobody can answer it: what documents do we actually have?

Not "what are our record series." Not "what does the retention schedule cover." The literal question: how many files, where do they sit, who owns them, and how many pages are in them. Agencies routinely discover their real number is five to ten times what staff estimated, which we have written about separately.

This article is about building the list.

Why the inventory comes first

An inventory is not busywork. Every downstream decision depends on it:

  • Retention. You cannot apply a retention schedule to records you have not located. Documents you forgot about are documents you are still legally holding.
  • Archiving. Deciding what moves to an archive section requires knowing what exists and when it was last touched.
  • Accessibility. Remediation is priced per page. Without a page count, every vendor quote you receive is a guess dressed up as a number.
  • Public records requests. The clock starts whether or not you know where things are.
  • Duplication. The same PDF linked from six department pages is one remediation job, not six — but only if you can see that it is the same file.

The seven places documents actually live

The single most common failure is inventorying one system and calling it done. Public documents at a typical city or district are spread across at least seven locations, and no two of them talk to each other.

LocationWhat is thereHow to enumerate it
Website CMSStaff reports, forms, notices, plans, brochuresCrawl the public site; export the media library
Agenda / meeting platformAgendas, packets, minutes, exhibitsPlatform export or crawl of the public portal
Records / ECM systemPermits, deeds, resolutions, scanned backfileSystem report by record series
Shared network drivesWorking files, drafts, decades of accumulationFile-system walk with size and modified date
Cloud collaborationRecent working documents, department wikisAdmin export by site or library
Department-run systemsUtility billing, permitting, GIS attachments, HRAsk each system owner for an attachment report
Physical storageBoxes, files, microfilm, plan drawersBox-level inventory; do not skip this

The last row is the one people leave out because it is not searchable. It is also usually the one holding the records with the longest retention obligations.

What to record for each document

Resist the urge to capture everything. An inventory with fifty fields never gets finished. These nine carry almost all the decision value:

  1. Stable identifier — URL, file path, or system ID
  2. File name and title — both, because they frequently disagree
  3. Page count — the unit that drives every cost estimate
  4. File size and format — flags scanned images and oversized files
  5. Content hash — SHA-256, which is how you find true duplicates
  6. Owning department — somebody has to make the keep-or-purge call
  7. Date created or published — drives retention and archive eligibility
  8. Last accessed or downloaded — the best available proxy for public demand
  9. Text layer present — yes or no; separates scanned images from real documents

On hashing: file names lie and URLs multiply, but a SHA-256 of the file bytes does not. Hashing is the only reliable way to prove that budget-fy24.pdf in the finance section and FY2024_Adopted.pdf in the clerk's archive are the same document. In our audits, content-level deduplication routinely removes 30–40% of a corpus.

Running the inventory in three passes

Pass 1: Machine count of everything public (1–5 days)

Crawl the public website and every public portal, following links to documents. Download each file, read its page count and text-layer status from the file itself, and hash it. This produces a complete, deduplicated list of what a member of the public — or a plaintiff's attorney — can reach today.

Do it with a crawler that follows document links (Screaming Frog and similar commercial tools, or wget plus pdfinfo if you have technical staff), or have a vendor run it. This pass is cheap and finishes fast.

Pass 2: System exports for everything non-public (1–3 weeks)

Ask each system owner for a report: every attachment, its size, its date, its record series if the system tracks one. The permitting system, the utility billing system, the HR system, the shared drives. You are not reading these documents yet. You are counting them and finding out who owns them.

Expect resistance and expect gaps. The goal is coverage, not perfection — a department that reports "roughly 40,000 scanned permit files in the ECM" has told you something useful even without a file list.

Pass 3: Physical and legacy media (ongoing)

Box-level is fine. Record location, series, date range, approximate linear feet or number of images, and condition. Microfilm and microfiche get counted in reels and rough image counts. This pass rarely finishes quickly; start it early and let it run alongside the others.

Traps that corrupt the count

  • Counting links instead of documents. One PDF linked from twelve pages is one document. Deduplicate by hash before you report a number.
  • Missing generated documents. Agenda platforms and permit systems often build PDFs on demand at URLs a crawler never sees. Ask the vendor for an export.
  • Login-gated content. Board portals and employee intranets hold documents that may still be public records.
  • Stale redirects. A 301 chain can hide a document behind three URLs and inflate your count threefold.
  • Counting documents, not pages. A 400-page board packet is one row in a spreadsheet and 400 units of remediation work.
  • Ignoring the archive that is already there. Many sites have a folder of old material nobody has looked at in years. It is still published, still public, and still counts.

What the inventory lets you do next

Once the list exists, four projects that previously felt impossible become straightforward scoping exercises:

  • Apply the retention schedule. Sort by series and date, identify what is eligible for disposition, and route it through your approval process. Some of your corpus will not need to be archived, scanned, or remediated at all, because it should be destroyed. See our guide to California retention schedules.
  • Decide what is archive-eligible. The ADA Title II archive exception can remove a meaningful share of your corpus from active remediation scope — if the content genuinely qualifies and is labeled correctly.
  • Prioritize scanning. Combine retention life with public demand to decide what to scan first.
  • Get real quotes. Hand vendors a page count and a duplicate rate instead of asking them to guess.

The order matters. Inventory, then retention, then archive decisions, then remediation. Agencies that remediate first routinely pay to make documents accessible that they were legally free to destroy, or that were duplicates of a file already fixed. That money does not come back.

We will build the inventory for you

A free corpus audit crawls every public document on your website, counts pages, flags duplicates, and hands you the spreadsheet. California public agencies, five business days.

Request a Demo More Articles