Back to blog

Purchase order extraction: PO numbers, parties, totals and delivery terms

Published , 4 min read

Short answer

For a purchase order, extract the PO number, order and delivery dates, buyer, supplier, order total, currency, payment terms and shipping method as single fields. Describe buyer and supplier by role so they are not swapped, and treat line items as out of scope for span extraction: ask for the total and the count instead.

A purchase order sheet with a highlighted table row beside two shipping boxes on a pallet

Which fields should a purchase order taxonomy have?

FieldTypeNotes
PO numberidHints: PO, purchase order
Order datedateReturned as an ISO date
Requested deliverydateHints: delivery, deliver by
BuyerstringThe purchasing company
SupplierstringThe company supplying the goods or services
Order totalmoneyThe total value of the order
CurrencyenumUSD, EUR, GBP, INR, SGD, other
Payment termsstringFor example, Net 30
Shipping methodstringCarrier or incoterms

The built-in preset adds buyer contact, number of line items and approver, for twelve fields in total.

How do you stop buyer and supplier being swapped?

On a purchase order the buyer is the issuer; on an invoice the supplier is. A layout rule such as "the name at the top" fails as soon as you mix the two document types. Describe each party by what it does ("The purchasing company", "The company supplying the goods or services") and add the labels documents use: "bill to" and "ship to" for the buyer, "vendor" for the supplier.

How are the two dates kept apart?

Both fields are typed date, so both are chosen from the same short list of dates on the page. The descriptions and hints separate them. When only one date is printed, the delivery field should come back absent; check that it does on your documents before trusting it.

What about line items?

Span extraction returns one value per field, so a table that should come back as an array of rows is outside what it does. Two things still work: the order total, which is a single amount, and a count of line items as a number field. If you need every row, pair this with a table extractor and use the fields here as the header record.

How do I run it?

curl -s https://jextract.com/api/extract \
  -F "file=@po.pdf" \
  -F "taxonomy=purchase_order" \
  -F "stream=0"

The response includes a flat data object keyed by field id, which is the part most integrations store, alongside the per-field detail used for review. Try it on the sample purchase order in the app.

Questions

Can it extract every line item on a purchase order?

No. The method returns one span per field, so tables that should come back as arrays are a known limitation. It returns header fields such as the PO number, parties, dates and total.

Does the same taxonomy work for invoices?

Use the invoice preset for invoices. The fields overlap, but the roles differ: an invoice is issued by the supplier and a purchase order by the buyer, and the descriptions reflect that.

How is the currency determined?

Currency is an enum field. It is judged from the document's context and returns one of the options you define, such as USD or EUR, with "other" as a fallback.

Run it on your own PDF.

Back to all posts