Scanned documents and OCR for accessibility
Why scanned PDFs fail every accessibility standard, how OCR fits into remediation, and how to decide between OCR, re-creation and archiving for a scanned backlog.
By DocoMatic Team
Published September 14, 2026 · Updated September 14, 2026
A scanned PDF is a photograph of paper: there is no text for a screen reader to read, no structure to navigate, nothing to search. Public archives are full of them — minutes, ordinances, old forms. This guide covers how OCR turns images back into text, where it fails, and how to triage a scanned backlog.
Why do scanned PDFs fail accessibility checks?
Because they contain no text. A scan is an image; assistive technology finds nothing to read, search finds nothing to match, and no amount of tagging fixes a page with no text layer. Every scanned document fails WCAG's text-alternative requirements until OCR or re-creation gives it real, correct text.
- Image-only PDFs vs. "searchable" scans with a hidden text layer
- Why tagging cannot start until a text layer exists
- How to detect scans in bulk across a document inventory
What does OCR actually do in a remediation pipeline?
OCR recognizes the characters in the page image and adds an invisible, selectable text layer aligned with the scan. In a remediation pipeline it runs first; tagging, reading order and metadata are built on top of the recognized text. OCR quality therefore caps the quality of everything after it.
- The OCR step: recognition, layout analysis, text layer placement
- What OCR does not do: structure, alt text, form fields
- Language settings and multi-language documents
When does OCR fail, and what then?
OCR degrades with the source: faxed copies, handwriting, stamps, low-resolution scans, tight multi-column layouts and tables. When accuracy drops below usable, the honest options are re-typing the document, recreating it from the original source file, or deciding it does not need to be online at all (Holley, D-Lib Magazine 2009).
- Failure modes: handwriting, degradation, complex layout, mixed languages
- Accuracy review: sampling recognized text against the image
- Re-creation vs. correction: when each is cheaper
Should old scanned documents be remediated at all?
Often no. The Title II archived-content exception can cover a scanned document that is kept only for reference or recordkeeping and is not currently needed to use the entity's services. Triage the scanned backlog first: a large share may be legitimately archived, retired, or held for on-request remediation.
- Applying the archived-content exception to scan archives
- On-request remediation as a policy for the long tail
- What must still be fixed: any scan people currently need
How does DocoMatic handle scanned documents?
DocoMatic detects image-only pages during discovery, runs OCR as a priced add-on (two credits per page), and flags low-confidence recognition for human review instead of publishing bad text. The output is a normal remediation candidate: real text, then tags, then verification like any born-digital file.
- Scan detection in the document inventory
- OCR confidence thresholds and human review routing
- Verifying OCR output before tagging begins
Sources
Related reading
Find the organizations this applies to
Every public organization we track, with its exact ADA Title II date; reports where we have crawled its website.
- ADA Title II deadlines by state
- California: deadlines for every public organization
- New York: deadlines for every public organization
- Texas: deadlines for every public organization
- Los Angeles County, California
Organizations with published reports
- California State University-Chancellors Office: deadline and document report
- Fort Bend County: deadline and document report
- Monroe County: deadline and document report
- California State University-Sacramento: deadline and document report
- San Diego State University: deadline and document report
- Denton County Transportation Authority: deadline and document report