OCR for scanned PDFs.
Scans, read. Not retyped. Turn English printed text into a searchable PDF or Markdown you can copy and use.
Read a scanned PDF
English printed text. Free, no account: 20 files a day and 200 pages in all, shared with the Markdown tool, up to 50 pages each. Uploads up to 4 MB; a PDF at a public link up to 50 MB.
Your file is processed on a worker we run on Modal and returned to your browser. OCR uses temporary files that our code removes after processing. Results get no public link. Hosting provider log retention has not been measured. Privacy and processing.
What to expect
A searchable PDF lets you find and select recognized words. Markdown gives you the recognized text. Neither is a human transcription: check names, numbers and tables before relying on them.
English printed text is supported at launch. Handwriting, other languages, faint copies and complex layouts can produce incomplete or incorrect text. Pages that already have text keep their text layer.
For automation, POST /api/v1/ocr uses the existing API allowance at one page credit per page. OCR API docs · API plans.
Measured agreement, including the misses
97.96% word agreement against the original PDF text layers on 50 controlled image-only scans. Measured 2026-10-01. This is a machine reference, not a human transcription or a guarantee of accuracy on your file.
We turned clean public-domain pages into image-only PDFs at 200 dpi, then compared the recognized words with their original text layers. Insertions, deletions, substitutions and reading order count against the score; case, punctuation and whitespace are normalized. This measures recognized text in the searchable PDF, not Markdown table structure.
The worst page scored 14.5% agreement: NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), page 16, a sideways diagram with poor OCR and fragile reference reading order. It stays in the reported total. Some visible figure words are missing from the original text layers, so correct OCR words can count as extra words against this reference.
These were clean raster scans, not physical scanner captures or phone photographs. The set does not measure handwriting, skew, damage, other languages or general OCR accuracy. Important names, numbers and tables still need checking against the source.