DocoMatic

Scanned documents and OCR for accessibility

Why scanned PDFs fail every accessibility standard, how OCR fits into remediation, and how to decide between OCR, re-creation and archiving for a scanned backlog.

By DocoMatic Team

Published September 14, 2026 · Updated September 14, 2026

A scanned PDF is a photograph of paper: there is no text for a screen reader to read, no structure to navigate, nothing to search. Public archives are full of them — minutes, ordinances, old forms. This guide covers how OCR turns images back into text, where it fails, and how to triage a scanned backlog.

Why do scanned PDFs fail accessibility checks?

Because they contain no text. A scan is an image; assistive technology finds nothing to read, search finds nothing to match, and no amount of tagging fixes a page with no text layer. Every scanned document fails WCAG's text-alternative requirements until OCR or re-creation gives it real, correct text.

  • Image-only PDFs vs. "searchable" scans with a hidden text layer
  • Why tagging cannot start until a text layer exists
  • How to detect scans in bulk across a document inventory

What does OCR actually do in a remediation pipeline?

OCR recognizes the characters in the page image and adds an invisible, selectable text layer aligned with the scan. In a remediation pipeline it runs first; tagging, reading order and metadata are built on top of the recognized text. OCR quality therefore caps the quality of everything after it.

  • The OCR step: recognition, layout analysis, text layer placement
  • What OCR does not do: structure, alt text, form fields
  • Language settings and multi-language documents

When does OCR fail, and what then?

OCR degrades with the source: faxed copies, handwriting, stamps, low-resolution scans, tight multi-column layouts and tables. When accuracy drops below usable, the honest options are re-typing the document, recreating it from the original source file, or deciding it does not need to be online at all (Holley, D-Lib Magazine 2009).

  • Failure modes: handwriting, degradation, complex layout, mixed languages
  • Accuracy review: sampling recognized text against the image
  • Re-creation vs. correction: when each is cheaper

Should old scanned documents be remediated at all?

Often no. The Title II archived-content exception can cover a scanned document that is kept only for reference or recordkeeping and is not currently needed to use the entity's services. Triage the scanned backlog first: a large share may be legitimately archived, retired, or held for on-request remediation.

  • Applying the archived-content exception to scan archives
  • On-request remediation as a policy for the long tail
  • What must still be fixed: any scan people currently need

How does DocoMatic handle scanned documents?

DocoMatic detects image-only pages during discovery, runs OCR as a priced add-on (two credits per page), and flags low-confidence recognition for human review instead of publishing bad text. The output is a normal remediation candidate: real text, then tags, then verification like any born-digital file.

  • Scan detection in the document inventory
  • OCR confidence thresholds and human review routing
  • Verifying OCR output before tagging begins

Sources

Related reading

Find the organizations this applies to

Every public organization we track, with its exact ADA Title II date; reports where we have crawled its website.

Organizations with published reports