Document Rescue Atlas / public evidence

Where ordinary OCR breaks.

Five public documents expose five different failure modes: handwriting, dense tables, faded notebooks, structured forms and right-to-left manuscript text. Inspect the source and the test status. These cases do not demonstrate five successful extractions.

5public source families
5distinct OCR failure modes
0customer files used
Human reviewrequired for every result

Inspectable cases

One rescue method does not fit every broken document.

The source images below illustrate document challenges. Earlier layout illustrations have been removed because they were not measured extraction results. The test status distinguishes observed failures from cases that remain unvalidated. Compare a printed source with its verified actual extraction.

SourceNIST SD19 handwritten sample form before document recovery

Case 01 / handwriting and form structure

NIST SD19 handwriting sample

Handwriting crosses printed guide lines and varies in spacing, stroke width and character shape. A useful recovery must preserve the form's field context instead of returning one undifferentiated block of text.

Test statusHandwriting is not a validated DocUnlocked use case. In our September 24, 2026 test of this NIST sample, none of four checked handwritten fields (date, city, state and ZIP code) appeared in the extracted text. No usable field table was produced. Do not purchase this service expecting reliable handwritten form data.

Source: NIST Special Database 19 sample image. Public benchmark source.

Review the handwritten form extraction workflow.

Source1930 US Census handwritten population schedule with dense rows and columns

Case 02 / handwriting inside a dense table

1930 Census population schedule

The text is only half the problem. A name attached to the wrong row or column becomes a confident-looking factual error. Recovery therefore needs row identity, column meaning and visible review flags.

Test statusThis handwritten 1930 Census page illustrates a difficult source, not a verified recovery. Its row and cell accuracy has not been validated with the current engine. The earlier layout illustration was not a measured extraction result.

Source: U.S. National Archives educational PDF.

Review the scanned table extraction workflow.

SourceAlexander Graham Bell laboratory notebook page with faded handwriting

Case 03 / faded notebook and reading order

Alexander Graham Bell notebook

Low contrast, historical handwriting and irregular line flow make plain OCR brittle. The useful outcome is a navigable note with uncertain passages exposed, not silently normalized prose.

Test statusOur September 24, 2026 test retained four of eight preselected text landmarks on this Bell page, with meaning-changing substitutions in other passages. This is not a full-page accuracy score. The page does not validate dependable handwriting transcription.

Source: Library of Congress, Alexander Graham Bell papers. See the source page for rights information.

Review the scanned PDF to Markdown workflow.

SourceIRS Form W-4 public form before structured document recovery

Case 04 / printed form and field semantics

IRS Form W-4

A clean-looking form can still fail when labels, checkboxes and values lose their relationships. Recovery quality depends on preserving field meaning and warning reviewers where a mark or value is ambiguous.

Test statusThis blank W-4 form illustrates field-layout challenges. It does not demonstrate completed-field recognition or automatic checkbox validation. Our separate printed-table test does not establish those capabilities.

Source: Internal Revenue Service public Form W-4 PDF. U.S. federal government document.

Review the unreadable PDF recovery workflow.

SourcePublic-domain Tashelhit manuscript page written in Arabic script

Case 05 / right-to-left historical script

Tashelhit manuscript in Arabic script

Historical glyph forms, connected script and right-to-left reading order create a different problem from Latin print. A responsible recovery keeps directionality and uncertainty visible for a qualified reviewer.

Test statusThis Arabic-script manuscript is an illustrative source only. No successful extraction or reading-order accuracy has been validated for this page.

Source: Wikimedia Commons public-domain manuscript image.

Review the old record search and provenance workflow.

OCR failure library

Diagnose the failure before choosing the recovery path.

A PDF can look readable to a person while its text layer is empty, scrambled or detached from the page structure.

Missing text layer

Search and copy return nothing because every page is only an image.

Reading-order collapse

Columns, notes and headers are extracted in the wrong sequence.

Table row drift

Correct words land under the wrong field or belong to the wrong record.

Handwriting ambiguity

Names, dates and numbers contain plausible but consequential substitutions.

Low-contrast loss

Faded ink disappears while paper texture becomes noise.

Mixed print and script

Printed labels survive while handwritten answers are dropped.

Directionality errors

Right-to-left lines or multilingual fragments are reordered incorrectly.

False confidence

The output looks fluent but hides unresolved passages and missing provenance.

Next decision

Check the file before paying for recovery.

The free browser check identifies format, size and basic recovery signals without uploading the document. If the file is accepted, the same flow can continue to the one-time $0.99 recovery.