Back to blog

PDF extraction API: from curl to structured JSON

Published , 5 min read

Short answer

Send one multipart POST to /api/extract with a PDF and a taxonomy. Add stream=0 for a single JSON response. You get a flat data object keyed by field id, plus per-field detail: the span, a normalised value, confidence, page, bounding box, status and the candidates it was chosen from.

A terminal window sending a request to a document, which returns a card with green curly braces

What is the smallest request that works?

curl -s https://jextract.com/api/extract \
  -F "file=@invoice.pdf" \
  -F "taxonomy=invoice" \
  -F "stream=0"

file is the PDF, up to 8 MB and 40 pages. taxonomy is a preset id: invoice, contract, resume or purchase_order.

Which options are there?

Form fieldValueEffect
fileA PDFThe document to read
taxonomyPreset id or JSONThe fields to extract
stream0Return one JSON object instead of server-sent events
ocr1Recognise scanned pages (slower)
refine0Skip the round that trims free-text spans

How do I send my own fields?

Pass a JSON object as the taxonomy field. Each field needs an id, name, type and description; enums also need options.

curl -s https://jextract.com/api/extract \
  -F "file=@policy.pdf" \
  -F "stream=0" \
  -F 'taxonomy={
    "name": "Insurance policy",
    "fields": [
      { "id": "policy_number", "name": "Policy number", "type": "id",
        "description": "The policy identifier" },
      { "id": "premium", "name": "Annual premium", "type": "money",
        "description": "Total premium per year" },
      { "id": "renews", "name": "Auto-renews", "type": "boolean",
        "description": "Whether the policy renews automatically" }
    ]
  }'

What comes back?

{
  "document": { "name": "invoice.pdf", "pages": 1, "chunks": 5 },
  "fields": [
    { "fieldId": "total", "text": "$4,149.39", "value": 4149.39,
      "confidence": 0.99, "page": 1,
      "bbox": { "x": 480, "y": 421, "width": 62, "height": 12 },
      "status": "found",
      "candidates": [ { "text": "$4,149.39", "p": 0.99 } ] }
  ],
  "data": { "invoice_number": "NW-2026-0912", "total": 4149.39, "currency": "USD" },
  "timing": { "parseMs": 4, "locateMs": 173, "pickMs": 158, "totalMs": 337 },
  "usage": { "requests": 2, "inputTokens": 9132, "costUsd": 0.00038 }
}
  • data is the part to store: one normalised value per field id. Dates are ISO strings, amounts are numbers.
  • fields is the part to review: the printed span, confidence, position and the alternatives.
  • status is found, low-confidence or absent. Absent fields have no value.
  • timing and usage report what the run took. The example above is a run with the refine round off.

When should I use streaming?

Leave stream unset when a person is watching. The response is then a server-sent event stream that reports each stage as it finishes: parsed, located, picked, refined, and finally done with the full result, or error with a message. For a backend job, stream=0 is simpler.

What should my code handle?

  • A 400 response when the taxonomy is missing or is not valid JSON with a fields array.
  • Fields with status absent: check the status before reading the value.
  • Fields with status low-confidence: route them to review with the bounding box.
  • Scanned files that return few or no fields: retry with ocr=1.

The API tab in the app has the same examples next to a live run.

Questions

What format does the API accept?

PDF files up to 8 MB and 40 pages, sent as multipart form data along with a taxonomy. Scanned PDFs need ocr=1.

Does the API return JSON or a stream?

Both. By default it streams server-sent events so an interface can show each stage. With stream=0 it returns the final result as one JSON object.

How do I get just the values?

Read the data object in the response. It is keyed by field id and holds one normalised value per field, without the review detail.

Run it on your own PDF.

Back to all posts