The table fails structurally
PSM 11 finds 531 word tokens in the Census schedule, but the plain-text output does not retain row identity or column attachment. More extracted tokens do not reconstruct a reliable census record.
Open OCR Failure Benchmark / v0.1
Eight reproducible Tesseract runs show what changes when the same public baseline meets a handwriting form, a census table, a faded notebook and a clean government form. Raw text, TSV, source hashes and calculations are public. Tesseract is the benchmark baseline, not the paid DocUnlocked recovery engine.




Measured findings
The percentages below are the share of recognized words below 50 in Tesseract's internal confidence signal. They are not character or word accuracy scores.
PSM 11 finds 531 word tokens in the Census schedule, but the plain-text output does not retain row identity or column attachment. More extracted tokens do not reconstruct a reliable census record.
NIST PSM 6 returns 67 words and emphasizes the handwritten passage. PSM 11 returns 218 words and surfaces more printed labels and number rows. A single baseline output hides this configuration sensitivity.
The W-4 PSM 11 run reports 90.04 mean confidence across 792 words. Its plain text still cannot represent checkbox state or guarantee label-to-field relationships.
The Bell outputs contain plausible English-like fragments while 85.3% to 87.4% of recognized words remain below 50 confidence. Review against the page image is part of the recovery task.
Complete metrics
| Case | PSM | Words | Characters | Lines | Mean confidence | Median | Below 50 |
|---|---|---|---|---|---|---|---|
| NIST SD19 | 6 | 67 | 414 | 14 | 45.81 | 41.21 | 55.22% |
| NIST SD19 | 11 | 218 | 1,418 | 72 | 65.15 | 83.88 | 33.94% |
| 1930 Census | 6 | 343 | 1,248 | 29 | 24.65 | 23.36 | 92.42% |
| 1930 Census | 11 | 531 | 2,394 | 355 | 29.32 | 29.17 | 83.99% |
| Bell notebook | 6 | 151 | 648 | 20 | 25.84 | 22.61 | 87.42% |
| Bell notebook | 11 | 143 | 577 | 47 | 26.31 | 23.42 | 85.31% |
| IRS W-4 | 6 | 784 | 4,458 | 55 | 86.75 | 95.25 | 7.14% |
| IRS W-4 | 11 | 792 | 4,685 | 103 | 90.04 | 95.70 | 3.41% |
Open data
Summary CSV · Summary JSON · Source manifest and hashes
Reproducible protocol
This release does not report character error rate or word error rate because a complete independently verified ground truth is not bundled for all four cases. It is a transparent failure-observation dataset, not a universal leaderboard.
Use the evidence