In many companies, someone still types data from PDFs and photos into another system all day. Supplier invoices into the accounting system, receipts into expense claims, customer ID cards and application forms into the CRM. It is slow, it is boring, and typing errors end up in reports and payments. Modern AI can do most of this work. The hard part is building a process that is accurate enough to trust with money and personal data. Here is how we approach it.
What Has Changed
Older OCR systems read text well but needed a template for every document layout. A new supplier with a different invoice design meant new configuration. Large language models with vision can read a document they have never seen before and find the invoice number, date, supplier, line items, and totals from context, the way a person would. This makes automation practical for the long tail of layouts that templates never covered.
How the Pipeline Works
- 1Intake: documents arrive by email, upload, scanner, or API, and each one gets an ID and is stored unchanged.
- 2Preparation: split multi-page PDFs, fix rotation, and detect the document type, such as invoice, receipt, or ID card.
- 3Reading: extract text with OCR or send the page image to a vision-capable model. Many systems do both and give the model the OCR text alongside the image.
- 4Extraction: ask the model for the fields you need in a fixed JSON schema, using the structured output feature of the model API so the response always matches it.
- 5Validation: check the result with ordinary code, such as whether line items add up to the total, whether the tax number has the right format, and whether the supplier exists in your master data.
- 6Review: send documents that fail validation or have low confidence to a person, with the document and extracted fields side by side.
- 7Export: post approved data to the accounting system, ERP, or CRM through its API, and keep a link back to the original file.
Validation Is Where Accuracy Comes From
A model that reads 97 percent of fields correctly still gets three fields wrong in every hundred, and it does not always know which ones. Business rules catch most of these. Totals must equal the sum of lines plus tax. Dates must fall in a sensible range. A purchase order number must exist and belong to the same supplier. Bank account details that differ from the supplier record should always go to review, since changed bank details are a classic sign of invoice fraud.
Common Document Types
| Document | Typical fields | Watch out for |
|---|---|---|
| Supplier invoices | Supplier, invoice number, dates, line items, tax, total, bank details | Duplicate invoices, changed bank accounts, totals that do not add up |
| Receipts for expense claims | Merchant, date, amount, category | Faded thermal paper, crumpled photos, handwritten amounts |
| ID cards and KYC documents | Name, ID number, date of birth, address | Personal data rules, image quality, and fraud checks that need specialised tools |
| Application and registration forms | Applicant details, choices, signatures | Handwriting, checkboxes, and fields left empty |
| Delivery notes and purchase orders | Reference numbers, items, quantities | Matching against the invoice and the order in your system |
Measure Before You Automate
- Collect a few hundred real documents that represent your suppliers and formats, including bad scans.
- Have people enter the correct values once to create a test set.
- Measure accuracy per field, since a wrong total matters more than a wrong address line.
- Measure the straight-through rate, meaning the share of documents that pass validation with no human touch.
- Rerun the test set whenever you change the prompt, the model, or the validation rules.
Start with every document reviewed by a person and the AI pre-filling the fields. As the measured accuracy for a document type stays high, let documents that pass all checks go straight through, and keep sampling some of them for review.
Personal Data and Security
ID cards, payslips, and application forms contain personal data. In Indonesia, the Personal Data Protection Law (UU PDP) sets obligations for how it is collected, processed, stored, and shared. Check where a model API provider processes and stores data, whether it is used for training, and how long it is kept. For sensitive documents, a model running on your own servers or in your own cloud account can keep data in your control. Encrypt stored files, limit who can open them, and delete them on a schedule that matches your retention policy.
The model reads the document. Your validation rules and review process decide whether the result can be trusted.
We build document extraction systems that connect to existing accounting, ERP, and CRM tools, starting with a pilot on your real documents so accuracy and savings are measured before a full rollout.
Key takeaways
- Vision-capable language models read new document layouts without templates, which makes automating data entry practical.
- Use structured output with a fixed schema, then validate with business rules in ordinary code.
- Send documents that fail checks or have changed bank details to human review.
- Measure field-level accuracy and straight-through rate on a test set of your own documents.
- Treat ID cards and forms as personal data, and choose model hosting that meets your data protection obligations.


