Back to blog

Contract data extraction: pulling parties, dates and terms from agreements

Published , 5 min read

Short answer

Contract fields that are a short span on the page extract well: party names, effective and end dates, fee amounts, notice periods in days, governing law and signatories. Fields that are a whole clause, such as a liability cap written as a sentence, are the weak spot of span-based extraction and should go to review.

A fanned multi-page contract with a signature and a fountain pen, three passages marked with green brackets

What can you extract from a contract?

The contract preset in jextract has fourteen fields. They fall into three groups, and the groups behave differently.

GroupFieldsHow it is found
Short spansAgreement title, parties, effective date, end date, fees, governing law, signatoriesChosen from typed candidates on the page
Numbers inside sentencesInitial term in months, payment due in days, termination notice in daysChosen from the numbers in the located clause
JudgmentsAuto-renews (yes or no)Judged from the clause, not picked as a span

How do you tell the two parties apart?

By describing their role, not their position. "The legal name of the first party (often the provider or vendor)" and "The legal name of the second party (often the customer or client)" work across layouts, and hints such as "licensor", "client" or "between" catch the wording different templates use. The same applies to signatories: one field per side, each tied to its party.

How are notice periods and terms extracted?

A notice period usually appears as "either party may terminate on sixty (60) days' written notice". Typing the field as number means the candidates are the numbers in the clause Jev located, and the description ("Notice period in days for termination for convenience") decides which one. Keep the unit in the field name so the value is unambiguous downstream.

Which contract fields are hard?

  • Clause-like values. A liability cap is often a sentence with conditions. A span-based method returns the nearest phrase, which may not carry the whole meaning.
  • Terms defined elsewhere. "The Fees set out in Schedule 2" points to another section; the located clause does not contain the amount.
  • Multi-column layouts where lines from two columns merge when the text is read.
  • Scanned, signed copies when OCR is off.

For these, treat the output as a pointer: the bounding box takes a reviewer to the clause, even when the span itself is not the full answer.

How do I run it on my own agreements?

Start from the preset and remove what you do not need. Fewer, sharper fields beat a long list of overlapping ones.

curl -s https://jextract.com/api/extract \
  -F "file=@msa.pdf" \
  -F "taxonomy=contract" \
  -F "stream=0"

Documents can be up to 8 MB and 40 pages. Open the app with the sample master services agreement to see the fields, boxes and candidates before wiring up the API.

Questions

Can it extract whole clauses from a contract?

Not reliably. The method returns spans chosen from candidates, which suits names, dates, amounts and numbers. Long clause-like values are a known limitation; use the returned bounding box to send a reviewer to the right place.

How long a contract can it handle?

Up to 40 pages and 8 MB per file. Long documents are split into windows of about 22k tokens, and each window is searched in parallel, so length adds little to the wall-clock time.

Does it detect auto-renewal?

Yes, as a boolean field. It is judged from the renewal clause and returns yes or no with a probability, instead of being picked as a span.

Run it on your own PDF.

Back to all posts