A workflow without a checking step ships its worst output

Volume changes what an error rate means
A tool that is right most of the time feels reliable when you use it a dozen times and watch each result. Run the same thing three hundred times a week and the occasional wrong output is not occasional any more, it is a regular event, and the only question is whether anyone sees it before a customer does.
What makes this different from ordinary software failure is the presentation. A broken script throws an error; a wrong answer here arrives formatted correctly, written confidently and indistinguishable in appearance from the right ones. Nothing about the output signals which category it is in.
So a workflow that has no checking step has effectively decided that its worst output is acceptable. That may be a reasonable decision for some tasks, but it should be a decision rather than an omission.
Asking the tool to check itself is a weak filter
A second pass reviewing the first does catch real errors, and it is cheap, so it is worth having. What it is not is verification. It will also invent faults in correct work, talk itself out of right answers, and miss errors that follow from the same misunderstanding that produced them, since the misunderstanding is shared.
It works best when the check is narrow and mechanical: does every claim here appear in the supplied source, are all the dates in the required format, does this list have the number of items requested. Broad requests to review quality produce a plausible review, which is not the same as a finding.
Framing helps. A check performed without sight of the reasoning that produced the answer is more useful than one that follows it in the same conversation, because agreement with an argument you have just seen is close to free.
The strongest checks are the ones that are not judgement
Anything that can be verified by a rule should be. Formats, ranges, totals, required fields, whether a quoted string genuinely appears in the source document, whether a referenced identifier exists in your own records. These catch a specific and important class of failure with complete reliability and no ambiguity.
This is the argument for structured intermediate output stated in a different way. A paragraph cannot be checked by a rule; a set of fields can. Designing the workflow so that the fragile parts arrive as data rather than as prose is what makes automatic checking possible at all.
Where a check cannot be mechanical, sampling is the honest fallback. Reading one in twenty carefully tells you the error rate and the failure shape, which is the information you need to decide whether the process is fit for its purpose. Reading none tells you nothing at all.
Put the check where the error is still cheap
A mistake caught at the extraction stage costs a re-run. The same mistake caught after publication costs a correction, a conversation and some proportion of a reader’s trust. Position of the check matters more than its thoroughness, and the instinct to add one final review at the end is usually the least efficient arrangement.
The general shape is to verify inputs early, verify the fragile intermediate result immediately after it is produced, and reserve the human pass for whatever is left. That way the expensive attention is spent on judgement rather than on catching format errors a rule would have caught for nothing.
It also matters that a check can fail loudly. A verification step whose failure produces a note nobody reads is decoration. Failures should stop the flow, or route the item to a person, or at minimum be counted somewhere that gets looked at.
Deciding how much checking a task deserves
The proportionate answer depends on two things: how expensive an error is, and whether the person receiving the output can detect one themselves. Internal drafting where the reader knows the subject needs very little. Anything going to a customer, into a record, or in front of someone with no way to evaluate it needs a great deal.
That second factor is the one usually left out. A wrong answer given to an expert is a nuisance and a wrong answer given to someone who trusts it is a different kind of event, and the same output can be both depending on where it lands.
Where the checking would cost as much as doing the work by hand, that is a genuine finding rather than a problem to engineer around. It means the task is not a good fit, and the useful response is to change what is being automated rather than to quietly drop the check.
Common questions
Is a second tool checking the first any better than self-checking?
Somewhat, because two systems fitted to different material fail in less correlated ways, so one may catch what the other missed. It is still two plausible-output generators agreeing or disagreeing, which is weaker evidence than it appears. Treat a disagreement as a flag worth a look rather than treating agreement as confirmation.
What proportion of output should a person review?
There is no general figure, and the useful approach is to start high, measure what you find, and reduce only if the error rate justifies it. Review intensity should also track change: after any alteration to the prompts, the tool or the input format, go back to reading everything for a while.
How do I stop a checking step from becoming a rubber stamp?
Give it something specific to look for and somewhere to record what it found. A reviewer asked to confirm that output looks fine will confirm it. A reviewer asked to check three named things and log the result will actually check them, and the log tells you whether the process is still working.
Features writer, Prompt After Prompt
Aarav covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.