The output that is nearly right is the one that gets through

Obvious failures are cheap
An output that is off-topic, absurdly formatted or plainly nonsensical costs almost nothing. It is spotted in seconds by whoever reads it next, regenerated, and forgotten. The visible failure rate of these tools is made of exactly this sort of thing, which is why the visible failure rate is a poor guide to the risk.
The expensive category is different in kind. It is an output that is well written, correctly structured, appropriate in tone, right about eleven things and wrong about the twelfth. Everything surrounding the error certifies it, and a reader who is checking for quality finds quality.
What makes this a distinct problem rather than a matter of degree is that the usual defences do not apply. Reading it again doesn’t help, because it reads well. Asking whether it looks right is answered in the affirmative, honestly, by the same surface that hides the error.
Where the near-misses concentrate
They cluster in the details that carry meaning without carrying attention. A negation dropped from a clause. A date that is the right format and the wrong year. The second of two similarly named entities. A condition attached to the wrong item in a list. A figure converted into different units.
They also cluster at the boundaries of an instruction. A summary that covers everything except the final section. A rewrite that applies a rule to the first three paragraphs. An extraction that handles every document except the ones in an unusual layout, silently.
And they cluster in transitions between steps. Material that passes from one stage to another loses whatever wasn’t carried in the format, and the loss is invisible downstream because the next step receives something that looks complete.
Reading is the wrong instrument for finding them
A subtle error survives reading because reading is a comprehension activity: it builds an understanding of what the text means, and a coherent text produces a coherent understanding regardless of whether the underlying facts hold. Nothing snags.
The checks that do work are the ones that compare against something. The source document, the previous version, a total that should reconcile, a rule that can be applied mechanically. Each of these tests a specific property rather than an overall impression, which is why they catch what reading misses.
Comparison against the source is the most useful and the most skipped, because it requires having the source open and looking item by item. It is dull work and there is no substitute for it where the stakes justify the effort.
A cheap partial version: check three specific claims chosen at random rather than reading the whole thing again. Three verified points tell you more about the reliability of a document than a second complete read, and they take less time.
The confidence problem is structural
Nothing in the output marks the sentence that was uncertain. A claim the process was sure of and a claim it constructed to fill a gap are written in the same register, at the same length, with the same fluency, and they arrive next to each other in the same paragraph.
This isn’t something better prompting removes, though asking for uncertainty to be marked does help a little and is worth doing where the format permits it. Treat any such markings as a hint about where to look first, not as a map of where the errors are.
The practical consequence is that reliability has to come from outside — from the design of the task, from mechanical validation, from a person who knows the subject. Assuming it comes from the fluency of the text is the underlying mistake in almost every incident of this kind.
Design so the near-miss cannot survive
The strongest response is structural rather than vigilant. Ask for output that contains a checkable trace: the quoted phrase a claim came from, the numbers a total was built out of, the section a summary drew on. Verification then becomes a comparison instead of an investigation.
Prefer tasks with a ground truth wherever there is a choice. A transformation of supplied material can be checked against the material; an unsourced assertion about the world cannot be checked against anything short of doing the research yourself.
And accept that some near-misses will get through anything you build, so the last question is what happens then. A process with a correction path, a version history and a named owner recovers. One that assumed correctness discovers the error from a customer, which is the same error at several times the cost.
Common questions
Why are subtle errors harder than obvious ones?
Because everything around them certifies them. An output that is well written, correctly structured and right about eleven things gives a reader every signal of quality, and reading again cannot help, since reading builds understanding rather than testing facts.
Where do these errors concentrate?
In details that carry meaning without attracting attention — dropped negations, plausible wrong dates, the second of two similar names, conditions attached to the wrong item, converted units — and at the boundaries of an instruction, such as a final section that was quietly skipped.
What catches them if reading does not?
Comparison against something external: the source document, a previous version, a total that must reconcile, a rule applied mechanically. Where a full comparison is too expensive, verifying three claims chosen at random tells you more than a second complete read.
Editor, Prompt After Prompt
Bhavna covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.