Pulling fields out of messy documents works, and the empty ones decide whether it is usable

Why extraction is a good fit
Turning a pile of inconsistent documents into rows and columns is genuinely hard by conventional means, because every supplier formats an invoice differently and every author writes a date differently. Rules that handle the observed cases break on the next batch, which is why this work has traditionally been done by people.
It suits a language tool well, because the task is recognition rather than reasoning. The answer is present in the source, the model doesn’t have to supply anything, and the output has a shape you can define exactly. Those three properties make it a rare case where the output is straightforwardly checkable.
The value shows up at volume. Extracting from five documents isn’t worth building anything for; extracting from five hundred a month, arriving in nine formats, is a real problem that this approach solves adequately when it is built with the right caution.
Define the fields as strictly as you can bear
The specification is the whole job. Each field needs a name, a type, a format and an explicit statement of what to do when the document doesn’t contain it. Dates in one written form. Amounts as numbers without currency symbols. Names exactly as they appear rather than tidied.
That last point matters more than it looks. A process that helpfully corrects a misspelled company name has destroyed the evidence that the source was inconsistent, and reconciling against the original later becomes impossible. Extraction should transcribe, and any normalising should happen in a separate, visible step.
Where a field can take one of a fixed set of values, list them and require one of the list. Free text in a field that should have been a category is the most common way an extraction pipeline produces output that nothing downstream can use.
Absent is a value and it must be available
The dangerous failure in extraction is not a misread figure. It is a field that was not present in the document and comes back filled with something plausible, because the shape of the task implies that every field gets an answer.
The instruction has to make absence legitimate and easy: a specific marker for not stated, and an explicit statement that guessing is worse than leaving it empty. This one addition changes the character of the output more than any other single thing, and it is the part people leave out.
It also produces useful information. A field that comes back empty in a third of documents tells you either that the sources genuinely do not carry it or that your definition does not match what they call it. Both are worth knowing, and neither is visible when the gaps are quietly filled.
Ask for the location too, where the format allows. A field accompanied by the line or phrase it came from can be spot-checked in seconds without opening the whole document, and that shortens verification enough to make it actually happen.
Validate mechanically before anything reads the output
Much of what can go wrong is catchable without judgement. Dates that are not dates. Amounts outside a plausible range. Totals that do not match their components. Identifiers that fail a format check. Required fields that are empty when the document type says they should not be.
These checks belong in ordinary code sitting after the extraction, and they are cheap to write. They turn a pile of results into two piles: the ones that passed every mechanical test and the ones a person needs to look at, which is a far better use of attention than reading everything.
The rejected pile is also the best guide to improving the specification. If a particular field fails validation repeatedly, the problem is nearly always the definition rather than the extraction, and the fix belongs in the field description.
Sample the passes, not just the failures
Mechanical validation only catches what it was told to look for. A date that is well formed and wrong passes every check, so a proportion of the passing results has to be read against the source by a person, chosen at random rather than by convenience.
How large a proportion depends on what the data feeds. Figures that will be paid, filed or published warrant a much heavier sample than a set of records used for rough internal analysis, and it is worth deciding that ratio explicitly rather than by how much time is left.
What that sample is measuring is a rate rather than a document. If one in fifty passing records is wrong, that number is a property of the process, and it should be known and stated to whoever uses the data rather than discovered by them later.
Common questions
Why is extraction more reliable than most tasks?
Because the answer is present in the source, so nothing has to be supplied. It is recognition rather than reasoning, and the output has a defined shape, which makes it one of the few cases where correctness can be checked mechanically against the document.
What is the most important instruction to include?
That a missing field must be marked as missing rather than filled. Without it, the shape of the task implies every field gets an answer, and plausible invented values are far more damaging than empty ones because nothing downstream can distinguish them.
How much of the output needs human checking?
Everything that fails mechanical validation, plus a random sample of what passes — a well-formed but wrong value passes every automatic test. The size of that sample should follow from what the data is used for, and the error rate it reveals should be told to whoever relies on it.
Consumer editor, Prompt After Prompt
Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.