Back to blog

Extracting data from scanned PDFs: when to turn OCR on

Published , 4 min read

Short answer

A text PDF already contains its characters and positions, so it parses in milliseconds. A scanned PDF is a picture of a page and needs OCR first. In jextract, OCR is off by default; turn it on for scans with the ocr=1 option. It runs locally with Tesseract, takes seconds instead of milliseconds, and is less accurate than a born-digital file.

A flatbed scanner with a skewed sheet on the glass and a clean page of text rising out of it

What is the difference between a text PDF and a scanned PDF?

Text PDFScanned PDF
Where it comes fromExported from softwareA scanner or phone camera
What it containsCharacters with positionsAn image of the page
Can you select text in a viewer?YesNo
Parsing in jextractMilliseconds, no OCRSeconds, needs OCR

The quickest test is to open the file and try to select a word. If you can, it is a text PDF and OCR adds nothing.

What does OCR change in the pipeline?

Only the first stage. LiteParse bundles Tesseract, so with OCR on, scanned pages are recognised on the same machine before lines and chunks are built. The locate and pick requests to Jev are the same either way: they receive chunks of text and do not know where the text came from.

What does OCR cost?

  • Time. Parsing goes from single-digit milliseconds to seconds per document.
  • Accuracy. Recognised characters can be wrong: a zero read as the letter O, a one as a lowercase L. A value that was misread on the way in cannot be picked correctly later.
  • Boxes. Bounding boxes come from the recognised words, so they are as good as the scan is straight.

How do I get better results from scans?

  • Get the original. If a text PDF of the same document exists, use it.
  • Scan straight and at a reasonable resolution. Skew and low resolution are the two biggest sources of misread characters.
  • Type fields narrowly. An id or money field is chosen from a short list of candidates, which limits the damage a stray character can do.
  • Lower the bar for review. Send more low-confidence fields to a person for scanned documents than for text ones.

How do I turn OCR on?

In the app, tick "OCR for scanned pages". In the API, add one form field:

curl -s https://jextract.com/api/extract \
  -F "file=@scan.pdf" \
  -F "taxonomy=invoice" \
  -F "ocr=1" \
  -F "stream=0"

Questions

Does jextract work on scanned documents?

Yes, with OCR enabled. LiteParse bundles Tesseract, so scanned pages are recognised locally. It is slower (seconds instead of milliseconds) and less accurate than born-digital PDFs.

Should I leave OCR on for everything?

No. On a text PDF it adds time and nothing else. Leave it off by default and turn it on for files where you cannot select text.

Is the scanned document sent to a third party for OCR?

No. OCR runs locally in the parser. Only the recognised text chunks are sent to Jev for the locate and pick questions.

Run it on your own PDF.

Back to all posts