PDF data extraction without generation: choosing spans instead of writing JSON
Published , 5 min read
Short answer
Extraction without generation means the model never writes a value. Code cuts candidate spans out of the document, and a model picks one of them with a probability. Every returned value is text that exists on the page, with the box it was read from, so it can be checked instead of trusted.

What goes wrong when an LLM writes the JSON?
The usual recipe is to paste the document into a prompt and ask for JSON that matches a schema. It works often enough to demo. The problems show up at volume.
- The value is new text. A total can come back rounded, a date reformatted, an identifier with one character changed, and nothing in the output says so.
- There is no location. To check a value someone has to find it in the PDF by eye.
- Output tokens are generated one after another, so latency grows with the number of fields.
- The JSON has to be parsed and validated, and a malformed response means a retry.
How does extraction as retrieval work?
jextract treats the document as a set of chunks and your taxonomy as a set of questions. It runs in four stages, and only two or three of them call a model.
- Parse. LiteParse reads the positioned text locally, in memory. Lines are rebuilt by baseline and grouped into chunks. No model is involved.
- Locate. One request to Jev asks, for every field at once, which chunk holds the value. "None" is always an option.
- Pick. Code cuts typed candidates out of the located chunk: dates, amounts, identifiers, label and value splits, proper nouns. A second request asks Jev which candidate is the value.
- Refine. For free-text fields only, a third request checks whether the span carries extra words and trims it to the exact ones.
Jev is a model from TypeSafe that answers questions with probabilities over the options it is given. It does not write text, which is the point: the answer to "what is the total?" can only be one of the amounts that are printed on the page.
What does a result look like?
Each field comes back with the span as printed, a normalised value, a confidence, the page, a bounding box, and the candidates it was chosen from.
{
"fieldId": "total",
"text": "$4,149.39",
"value": 4149.39,
"confidence": 0.99,
"page": 1,
"bbox": { "x": 480, "y": 448, "width": 48.9, "height": 12.3 },
"status": "found",
"candidates": [
{ "text": "$4,149.39", "p": 0.99 },
{ "text": "(none)", "p": 0.01 }
]
}How fast is it, and what does it cost?
On the sample one-page invoice with a twelve-field taxonomy, one captured run took 813 ms end to end: 3 ms to parse, 475 ms to locate, 161 ms to pick and 173 ms to refine. It used three requests, 13,150 input tokens and no generated tokens, for $0.00055. Because both main requests ask about every field in parallel, adding fields barely changes the time. The architecture page shows that run request by request.
Where does this approach fail?
A method that only returns spans cannot return things that are not spans. It is the wrong tool for these cases:
- Long clause-like values, such as a liability cap written as a sentence.
- Fields whose description is ambiguous, such as a current job title against a headline title.
- Multi-column prose where lines merge across columns.
- Tables that should come back as arrays of rows.
- Scanned pages when OCR is turned off.
The candidates list on every field shows what the model was offered, so these cases are easy to spot: the right answer is simply not among the options.
How do I try it?
Open the app, pick a sample document or drop your own PDF, and run it. Or call the API with a preset taxonomy:
curl -s https://jextract.com/api/extract \ -F "file=@invoice.pdf" \ -F "taxonomy=invoice" \ -F "stream=0"
Questions
Can jextract return a value that is not in the document?
No. Values are chosen from spans that code cut out of the page, so the returned text always exists in the document. Enum and boolean fields are the exception by design: they are judged from context and return one of the options you defined.
Does extraction without generation still use an LLM?
It uses TypeSafe's Jev model, which answers multiple-choice questions with a probability for each option. It produces no output tokens, so there is no JSON to parse and no decode step that grows with the number of fields.
How many model requests does one document take?
Two, plus an optional third. One request locates every field, one picks every value, and a refine request runs only for free-text fields whose span might carry extra words. Long documents are split into windows of about 22k tokens that are located in parallel.