Blog

Guides to getting structured data out of documents in a form you can check: the method, the schema for each document type, and the numbers that come back.

A sheet of paper with boxed phrases, and tweezers lifting one green box off the page
, 5 min read

PDF data extraction without generation: choosing spans instead of writing JSON

Most PDF extractors ask an LLM to write JSON. jextract asks a model to choose a span that already exists on the page. Here is how that works and what it changes.