Extracting fields from long documents: chunks, windows and parallel search
Published , 5 min read
Short answer
jextract splits a document into chunks of nearby lines, groups the chunks into windows of about 22k tokens, and asks Jev to locate every field in each window in parallel. The per-window answers are merged, and the value is then picked once from the best chunk. Length adds requests that run side by side, not time in sequence.

Why are long documents harder?
A one-page invoice has a handful of places a value can be. A forty-page agreement has hundreds. Reading the whole thing in one go costs more, and the more text a question is asked against, the easier it is for a look-alike value on the wrong page to win.
How is a document split into chunks?
By layout, not by a fixed character count. Lines are rebuilt from positioned text, and a new chunk starts where the vertical gap between lines is large, or when a chunk reaches about 900 characters or 14 lines. The result follows the document's own blocks: a header, an address block, a table, a totals box. Each chunk keeps its page and bounding box.
What is a window?
A group of consecutive chunks that fits in one request, about 22k tokens. A short document is a single window. A long one becomes several, and each gets its own locate request.
- Every window is asked the same questions at the same time: for each field, which chunk here holds the value, or none.
- Each window answers with a probability for each of its chunks and for "none".
- The answers are merged, with each window weighted by how much probability it did not put on "none". A window that is sure the field is elsewhere barely counts.
- The pick request runs once, on the best chunk and its runner-up or neighbour.
What does length cost in time and money?
Locate requests run in parallel, so wall-clock time grows slowly with length. Cost grows with the input tokens read, which is roughly the size of the document, once. There are no output tokens to pay for, because nothing is generated.
What are the limits?
- Files up to 8 MB and 40 pages per request.
- Values that span chunks. A definition on page 2 and the amount it refers to on page 30 are not joined.
- Repeated values. The same field may be stated in a summary and again in a schedule; the description should say which one you mean.
How can I see what happened on my document?
After a run, the Jev calls view in the app lists every request with the chunks it was sent and the probability on each option. The architecture page shows the same view for a captured run.
Questions
How many pages can one request handle?
Up to 40 pages and 8 MB. Longer documents should be split before upload.
Does a longer document mean more model requests?
Yes for the locate step: one request per window of about 22k tokens, all run in parallel. The pick step still runs once.
Why not put the whole document in one prompt?
Jev chooses among options, and offering every span of a long document as an option would be enormous. Locating the chunk first keeps the pick question to a short list of typed candidates, which also keeps the probabilities sharp.