Skip to content
Working with the tools, not about them

Workflows

Reviewing two hundred outputs is a different skill from reviewing one

Checking at volume fails through attention rather than through ignorance, so the review has to be designed around how a person actually degrades over an hour.

By Bhavna Deshpande3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The reviewer is the bottleneck, and the reviewer gets tired

Generating two hundred items is quick. Reading two hundred items carefully is several hours, and the last fifty get a fraction of the attention the first fifty received. This isn’t a discipline problem; it is what sustained comparison of similar material does to anyone.

The characteristic pattern is worth recognising. Attention is high for the first dozen, settles into a rhythm, and then converts into pattern-matching. By the middle of the batch the reviewer is confirming that each item looks like the others rather than checking whether it is correct, and those are entirely different activities.

Fluent, uniform output makes this worse, because there is nothing to snag on. A batch of documents written by twenty different people is tiring in a way that keeps you alert. A batch written in one voice is soporific.

Sort before you read

The single most useful move is to stop reading in the order the items were produced. Sort by something meaningful first: length, whether a required field is empty, whether a validation failed, how unusual the source was, how much money or risk is attached.

Sorting concentrates the errors. Outputs that are much shorter or much longer than the rest are disproportionately the broken ones, and reading the extremes first finds most of the serious problems in a small fraction of the time.

It also lets you stop early with a defensible reason. Having read every unusual item and a sample of the ordinary ones, you know something about the batch. Having read the first sixty in arrival order, you know about the first sixty.

Check one property at a time across the whole batch

Reading each item completely, and judging everything about it at once, is slow and unreliable. Passing through the whole batch looking only at one thing — are all the dates present, does every piece address the right audience, is any figure outside range — is faster and catches far more.

This works because a single criterion holds in the mind cleanly, so the check becomes almost mechanical. Five passes over a batch, each looking for one thing, take less total time than one pass looking for everything and produce a much better result.

It also makes the review reportable. At the end you can say which properties were checked across all items and which were only sampled, and that is a genuinely useful statement to give to whoever is relying on the output.

Order the passes by consequence. The property whose failure would do the most damage gets checked across everything; the cosmetic ones get a sample, or get skipped with the decision recorded.

Look for what is the same, not only for what is wrong

At volume, the interesting failures are systematic rather than individual. Every item making the same unwarranted assumption, every summary omitting the same section, every response ending with the same inapplicable recommendation. These are invisible when items are judged one at a time and obvious when they are placed side by side.

So build in a comparison step: put ten outputs in one view and read across them rather than down. Repetition that would pass unnoticed in a single item becomes glaring, and one systematic fault found this way is worth more than twenty individual corrections.

A systematic fault also has a systematic fix. Individual corrections leave the process producing the same error next week, which is how a review turns into a permanent job instead of a temporary one.

Decide the disposition before you start

A review needs a rule for what happens to a flagged item. Fix it here, send it back for regeneration, drop it, escalate it. Deciding that during the review is what turns two hours into four, because each borderline case becomes a small negotiation with yourself.

The same applies to the standard. Write down what would make an item unacceptable before opening the first one, because standards drift downwards over a long batch, and they drift fastest when the alternative to accepting something is doing it yourself.

And if a review consistently finds that a third of a batch needs fixing, the honest conclusion is about the process rather than the batch. At that rate, checking and repairing is most of the work, and whatever was supposed to be saved by generating at volume has already been spent.

Common questions

Why does reviewing a long batch fail?

Because sustained comparison of similar material turns checking into pattern-matching. By the middle of a batch the reviewer is confirming that items resemble each other rather than testing whether they are right, and uniform fluent output accelerates that.

What order should a batch be read in?

Not the order it was produced. Sort by length, by failed validations, by unusual sources or by attached risk, and read the extremes first — outputs that are much shorter or longer than the rest are disproportionately the broken ones.

How do I catch faults that affect every item?

By comparing rather than reading. Place ten outputs side by side and read across them; a shared unwarranted assumption or a repeated omission is invisible item by item and obvious in comparison. Systematic faults also have systematic fixes.

Workflowsreviewquality controlbatchattention
Bhavna Deshpande
Editor, Prompt After Prompt

Bhavna covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.