Document Rescue Atlas / public evidence

Where ordinary OCR breaks.

Five public documents expose five different failure modes: handwriting, dense tables, faded notebooks, structured forms and right-to-left manuscript text. Inspect the source, the recovery preview and the uncertainty before trusting any output.

5public source families
5distinct OCR failure modes
0customer files used
Human reviewrequired for every result

Inspectable cases

One rescue method does not fit every broken document.

These are visual recovery previews, not accuracy certificates. Each case records what makes the source difficult, what becomes reviewable and what still needs a human decision.

SourceNIST SD19 handwritten sample form before document recovery
Recovery previewReview-ready text and fields recovered from the NIST handwriting sample

Case 01 / handwriting and form structure

NIST SD19 handwriting sample

Handwriting crosses printed guide lines and varies in spacing, stroke width and character shape. A useful recovery must preserve the form's field context instead of returning one undifferentiated block of text.

Recovered valueReadable field-oriented text for review and reuse.
Known limitNames, numbers and ambiguous letterforms still require comparison with the source.

Source: NIST Special Database 19 sample image. Public benchmark source.

Review the handwritten form extraction workflow.

Source1930 US Census handwritten population schedule with dense rows and columns
Recovery previewStructured review fields recovered from the 1930 Census schedule

Case 02 / handwriting inside a dense table

1930 Census population schedule

The text is only half the problem. A name attached to the wrong row or column becomes a confident-looking factual error. Recovery therefore needs row identity, column meaning and visible review flags.

Recovered valueSearchable names and fields organized for row-by-row checking.
Known limitDitto marks, faint pencil and column drift can change family relationships.

Source: U.S. National Archives educational PDF.

Review the scanned table extraction workflow.

SourceAlexander Graham Bell laboratory notebook page with faded handwriting
Recovery previewReview-ready notes recovered from the Bell laboratory notebook page

Case 03 / faded notebook and reading order

Alexander Graham Bell notebook

Low contrast, historical handwriting and irregular line flow make plain OCR brittle. The useful outcome is a navigable note with uncertain passages exposed, not silently normalized prose.

Recovered valueSearchable notes with a review path back to the page image.
Known limitScientific notation, insertions and faint words may remain unresolved.

Source: Library of Congress, Alexander Graham Bell papers. See the source page for rights information.

Review the scanned PDF to Markdown workflow.

SourceIRS Form W-4 public form before structured document recovery
Recovery previewStructured fields and review flags recovered from IRS Form W-4

Case 04 / printed form and field semantics

IRS Form W-4

A clean-looking form can still fail when labels, checkboxes and values lose their relationships. Recovery quality depends on preserving field meaning and warning reviewers where a mark or value is ambiguous.

Recovered valueReview-ready labels, fields and structural cues.
Known limitRecovered content is not tax advice and must be checked against the original form.

Source: Internal Revenue Service public Form W-4 PDF. U.S. federal government document.

Review the unreadable PDF recovery workflow.

SourcePublic-domain Tashelhit manuscript page written in Arabic script
Recovery previewSearchable review package preview for the Arabic-script manuscript

Case 05 / right-to-left historical script

Tashelhit manuscript in Arabic script

Historical glyph forms, connected script and right-to-left reading order create a different problem from Latin print. A responsible recovery keeps directionality and uncertainty visible for a qualified reviewer.

Recovered valueSearchable material organized into a review package.
Known limitLanguage expertise is essential before treating the transcription as authoritative.

Source: Wikimedia Commons public-domain manuscript image.

Review the old record search and provenance workflow.

OCR failure library

Diagnose the failure before choosing the recovery path.

A PDF can look readable to a person while its text layer is empty, scrambled or detached from the page structure.

Missing text layer

Search and copy return nothing because every page is only an image.

Reading-order collapse

Columns, notes and headers are extracted in the wrong sequence.

Table row drift

Correct words land under the wrong field or belong to the wrong record.

Handwriting ambiguity

Names, dates and numbers contain plausible but consequential substitutions.

Low-contrast loss

Faded ink disappears while paper texture becomes noise.

Mixed print and script

Printed labels survive while handwritten answers are dropped.

Directionality errors

Right-to-left lines or multilingual fragments are reordered incorrectly.

False confidence

The output looks fluent but hides unresolved passages and missing provenance.

Next decision

Check the file before paying for recovery.

The free browser check identifies format, size and basic recovery signals without uploading the document. If the file is accepted, the same flow can continue to the one-time $9 recovery.