Message match
You cannot select or copy the text.
Many scanned PDFs contain page images instead of a usable text layer. The document may open and look readable, but selecting, copying or searching returns nothing useful. DocUnlocked makes a new extraction attempt from the visible page content and prepares editable files for review.
Get editable files, not a rebuilt PDF promise.
The review package includes DOCX, HTML, Markdown, TXT and a generic document.json representation, together with summary, source-manifest and review-flag files. It does not promise a new searchable PDF, preservation of the original page layout or exact reconstruction of columns and visual styling.
When your first OCR output is unusable.
A first OCR pass can flatten columns, merge lines or return plausible-looking character substitutions. This workflow makes another extraction attempt with structure and source references, but it does not guarantee repair. Names, dates, identifiers, amounts and flagged passages need visual comparison with the scan.
Tables are conditional.
CSV and XLSX files are included only when detectable table structure supports them. They are absent when rows and columns cannot be identified reliably. Any extracted table still requires cell-by-cell review; formulas and spreadsheet logic are not reconstructed.
Questions before uploading.
Will you make my original PDF searchable?
DocUnlocked extracts the document into editable, review-ready files. It does not guarantee a rebuilt searchable PDF or preservation of the original page layout.
Is the text guaranteed accurate?
No. Accuracy depends on scan quality, fonts, handwriting and layout. Review the files and low-confidence notes before use.
Do tables always become CSV or Excel?
No. CSV and XLSX files are included only when detectable table structure supports them.
Is document.json a custom schema?
No. It is a generic document representation, not a custom business schema or API response.