OCR a scan to text (English)
Extract text from a scanned PDF with offline OCR. English recognition only, output is a plain .txt file — here's exactly what to expect.
The OCR tool in 1FileTool turns a scanned PDF into a plain text file on your Mac — no upload, no cloud OCR service. It recognises English text only and writes a .txt file, not a searchable PDF.
Protect your originals first
Replace source is ON by default: the output overwrites the original file and no backup is kept. Before following along, open Settings › General and turn Replace source off, or pick a separate output folder.
Run OCR on a scan
- Open PDF Tools › OCR (tool page).
- Drop the scanned PDF.
- Run it. Each page is rendered at 300 DPI and recognised, so a long scan takes a while — progress is per page.
- Find
<name>_ocr.txtin your output location. Pages are separated by a--- Page Break ---line.
What you get (and what you don't)
- The output is a plain
.txt— words and line breaks, no layout, no fonts. A two-column scan reads straight through. - Recognition is English-only. Pages in other languages come out garbled or empty.
- Clean, flat scans at reasonable contrast recognise well. Handwriting, stamps, skewed photos and faint thermal-paper text will need manual cleanup.
- OCR does not produce a searchable PDF — it extracts text. To get editable paragraphs in Word instead, use Scanned PDF to Word.
If you need the text inside another format
- Scanned PDF to Word runs the same OCR but writes a
.docxof plain paragraphs — see PDF to Word. - PDF Tools › Extract Text copies the embedded text layer of a digital PDF — instant, but it finds nothing on a scan because a scan has no text layer.
Related
- PDF to Word: visual vs text mode
- Protect and redact PDFs — OCR first if Detect PII reports a scanned file
- Local OCR and privacy
- Why files never upload
- Compress PDFs
Tools used