Extracting data from scanned PDFs: when to turn OCR on
Published , 4 min read
Short answer
A text PDF already contains its characters and positions, so it parses in milliseconds. A scanned PDF is a picture of a page and needs OCR first. In jextract, OCR is off by default; turn it on for scans with the ocr=1 option. It runs locally with Tesseract, takes seconds instead of milliseconds, and is less accurate than a born-digital file.

What is the difference between a text PDF and a scanned PDF?
| Text PDF | Scanned PDF | |
|---|---|---|
| Where it comes from | Exported from software | A scanner or phone camera |
| What it contains | Characters with positions | An image of the page |
| Can you select text in a viewer? | Yes | No |
| Parsing in jextract | Milliseconds, no OCR | Seconds, needs OCR |
The quickest test is to open the file and try to select a word. If you can, it is a text PDF and OCR adds nothing.
What does OCR change in the pipeline?
Only the first stage. LiteParse bundles Tesseract, so with OCR on, scanned pages are recognised on the same machine before lines and chunks are built. The locate and pick requests to Jev are the same either way: they receive chunks of text and do not know where the text came from.
What does OCR cost?
- Time. Parsing goes from single-digit milliseconds to seconds per document.
- Accuracy. Recognised characters can be wrong: a zero read as the letter O, a one as a lowercase L. A value that was misread on the way in cannot be picked correctly later.
- Boxes. Bounding boxes come from the recognised words, so they are as good as the scan is straight.
How do I get better results from scans?
- Get the original. If a text PDF of the same document exists, use it.
- Scan straight and at a reasonable resolution. Skew and low resolution are the two biggest sources of misread characters.
- Type fields narrowly. An
idormoneyfield is chosen from a short list of candidates, which limits the damage a stray character can do. - Lower the bar for review. Send more low-confidence fields to a person for scanned documents than for text ones.
How do I turn OCR on?
In the app, tick "OCR for scanned pages". In the API, add one form field:
curl -s https://jextract.com/api/extract \ -F "file=@scan.pdf" \ -F "taxonomy=invoice" \ -F "ocr=1" \ -F "stream=0"
Questions
Does jextract work on scanned documents?
Yes, with OCR enabled. LiteParse bundles Tesseract, so scanned pages are recognised locally. It is slower (seconds instead of milliseconds) and less accurate than born-digital PDFs.
Should I leave OCR on for everything?
No. On a text PDF it adds time and nothing else. Leave it off by default and turn it on for files where you cannot select text.
Is the scanned document sent to a third party for OCR?
No. OCR runs locally in the parser. Only the recognised text chunks are sent to Jev for the locate and pick questions.