Fundamentals · Document Processing & Validation
Document Extraction Workbench
From Documents to Validated Records
I will build an invoice extractor that keeps both the value and the evidence for every field. It will read text and scanned documents, validate each record against explicit rules, and send anything uncertain to a person.
Status: In Progress
The architecture and build plan are complete, and I am building this alongside Ask TanStack Query. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens.
The Problem
Picture a finance team that copies invoice details into a spreadsheet by hand: vendor, invoice number, date, currency, and total. It is slow, and a single wrong amount is expensive.
An AI model can read the invoice, but a confident wrong answer is worse than a blank one. The team needs a draft record they can check quickly, with the source of every value one click away. The goal is a reviewable draft, not payment approval.
How It Will Work
Reading a file, recognizing text, extracting meaning, checking a record, and measuring accuracy are separate operations, and the pipeline will keep them separate:
What I Will Build
- A typed record: Pydantic models for vendor, invoice number, date, currency, and total, each with a quote and page reference
- A reader: Page-by-page text extraction that detects scanned pages instead of silently returning empty text
- A validator: Normalization, quote checks, and review flags, with the reasons recorded
- An evaluator: A labeled dataset and a per-field accuracy report
The first version will run on saved example predictions so the whole pipeline can be tested before any paid model call. Live extraction with a real model comes after that.
Planned Design Decisions
Every value carries its evidence
Each field will be stored as a value, an exact quote, and a page number, so a reviewer can check the source in seconds.
Missing means null
A field that is absent or ambiguous will be null, never guessed. A currency symbol alone will not be treated as a currency.
Extraction and normalization are separate
Raw model output will be preserved next to the normalized record, so any change can be traced. Money will use Decimal, not floating point.
Business rules live in code
Date, currency, and amount rules will be explicit and tested, not left to a sentence in a prompt.
Scans take a separate route
A PDF parser cannot read an image. Scanned pages will use a visual route and will always require human verification of the evidence.
Documents are untrusted input
The prompt will tell the model to ignore instructions written inside an invoice. That is a defense, not a guarantee, so validation still runs afterward.
How I Will Evaluate It
- •A fresh set of 20 synthetic invoices: 5 clean, 5 with a missing or ambiguous field, 5 that separate subtotal, tax, discount, and final total, and 5 scans or multi-page documents
- •Ten documents for development and ten held out before any tuning
- •Labels written and checked by hand, never produced by the same model that makes the predictions
- •Exact-match accuracy by field, incorrect non-null values, correctly missing fields, review rate, and latency, with mock and live results reported separately
- •A deliberate failure case: a total read as 9.00 instead of 90.00 passes format checks and a quote check, and only the comparison with labeled truth catches it
A low review rate is not the goal on its own: a system can avoid mistakes by sending every document to a person. I will report the tradeoff between errors caught and review effort, and I will publish results only after they exist.
Scope
The workbench will produce draft records for human inspection. It will not approve or pay invoices, and a record without review flags will mean only that these particular checks found no issue, not that the invoice is proven correct. It will use synthetic invoices and support USD, EUR, and GBP, with credit notes out of scope.
