Document recovery guide

Recover a scanned table as rows and columns you can review.

Table extraction is not ordinary OCR and it is not automatic spreadsheet reconstruction. The system must detect a grid, attach each value to the correct row and column, and keep ambiguous cells visible before any export.

Public source and recovery preview

1930 Census Orchard Street schedule, U.S. National Archives

Detection before export

  1. Deskew and crop the page image.
  2. Detect grid lines or repeated column alignment.
  3. Segment cells and recover their text.
  4. Review merged, faint or handwritten cells.
  5. Keep CSV and XLSX absent unless a table is actually detected.

Table OCR versus spreadsheet reconstruction

OCR recognizes symbols. Table extraction adds row and column relationships. Spreadsheet reconstruction goes further by inferring formulas, types and business meaning. DocUnlocked provides reviewable extracted values; it does not recreate an authoritative workbook or guarantee formulas.

Ambiguity must remain visible

Merged cells, ditto marks, multi-line values and faint gridlines should produce review flags rather than silent shifts. A single displaced cell can corrupt every downstream row, so visual comparison is required before analysis.

Limits and failure modes

Questions about this workflow

Are CSV and XLSX always included?

No. They are conditional outputs produced only when table structure is detected.

Does this rebuild formulas or spreadsheet logic?

No. It extracts reviewable rows and columns; formulas and business rules must be recreated separately.

Which cells should I review first?

Names, numeric identifiers, merged cells, multi-line entries and any value marked uncertain.

Next best step

Run the free readability check before uploading. Working with a historical record? Return to the DocUnlocked hub.

Recover one document - $9