What a confidence score means in document extraction
Published , 5 min read
Short answer
In jextract, a field's confidence is the probability the model gave to the chosen span, multiplied by the probability that the field is present in the document at all. Fields below 0.55 are marked low-confidence, and a field is reported absent when "none" wins or reaches 0.6. The number comes from the model's answer distribution, not from asking it how sure it feels.

Why does the confidence number matter?
Nobody reviews every field of every document. A confidence score is what decides which fields a person looks at. If the number is made up, the review queue is either everything or nothing.
Where does the number come from?
Every question jextract asks is multiple choice, and the model returns a probability for each option. Two of those probabilities make up the score.
- Locate asks which chunk holds the field. One option is always "none". The probability on "none" is the chance the field is not in the document.
- Pick asks which candidate span is the value. The probability on the chosen candidate is how clearly it beat the others.
- Confidence is the pick probability multiplied by one minus the "none" probability.
confidence = p(chosen span) × (1 − p(none))
Enum and boolean fields are judged from context instead of picked from spans, so their confidence is the choice probability alone. When the refine round trims a free-text span, confidence is scaled by how sure the model was of the trimmed window.
What does that look like on a real document?
These are four fields from one run on the sample invoice, with the candidates the model was offered.
| Field | Chosen span | Confidence | Runner-up |
|---|---|---|---|
| Total due | $4,149.39 | 99% | none, 1% |
| Invoice date | September 12, 2026 | 94% | none, 6% |
| Due date | October 12, 2026 | 81% | none, 19% |
| IBAN / account | GB29 NWBK 6016 1331 9268 19 | 58% | none, 22% |
The total has one obvious candidate, so nearly all the probability lands on it. The IBAN is correct but scored 58%, because a second candidate that included the "IBAN:" label took 19% and "none" took 22%. That is the score doing its job: the value is right, and the number tells you the choice was closer than the others.
What do found, low-confidence and absent mean?
| Status | When | What to do |
|---|---|---|
found | Confidence is 0.55 or higher | Use it; sample for audit |
low-confidence | A value was chosen but confidence is under 0.55 | Send to review with the box and candidates |
absent | "None" won, or its probability reached 0.6 | Treat as missing, not as an error |
Absent is a result, not a failure. An invoice without a purchase order number should say so, and a system that always returns something for every field will invent one.
How should I use confidence in a workflow?
- Set the review threshold per field, not per document. A total deserves a stricter bar than payment terms.
- Show the reviewer the bounding box. Checking a highlighted span takes a second; finding it takes a minute.
- Look at the candidates before changing a threshold. A low score with the right answer in second place is a description problem, not a model problem.
- Measure on your own documents. The thresholds above are jextract's defaults, and the right bar for your data is the one that matches your error rate in review.
You can see the candidates and probabilities for every field in the app, and the exact requests behind them on the architecture page.
Questions
Is a 0.9 confidence a 90% chance of being correct?
It is the model's probability for that span, adjusted for the chance the field is absent. It is designed to track correctness, but you should verify it on your own documents before relying on a threshold.
Why is a correct value sometimes low confidence?
Because another candidate was a close alternative. A span with its label attached, a second date on the page, or a real chance that the field is missing all take probability away from the winner.
What happens when a field is not in the document?
It is returned with status absent and no value. "None" is an option in every locate question, so a missing field is reported as missing instead of being filled with the nearest look-alike.