Run twenty before you run two thousand

Scale converts a small flaw into an expensive one
A process that works on the three examples you tried will be applied to two thousand items, and every property of it gets multiplied: the cost, the time, the error rate, and the effort of cleaning up afterwards. A tendency that was a curiosity at three items is a project at two thousand.
Reversal is the part people underestimate. Records written into a live system, files renamed, messages sent, entries published — undoing these is often harder than doing them, and sometimes impossible. The pilot is worth running mainly because it happens before that door closes.
The other multiplied quantity is your own attention. A batch that requires a person to check every result has not saved anything, and whether that is the case is exactly what a pilot tells you, before the schedule has been set and the expectation created.
Choose the sample deliberately, not by convenience
Taking the first twenty items is the common approach and a poor one, because the first items are usually the most recent, the best formatted or the most typical. A sample chosen that way tells you how the process performs on easy material, which was never the open question.
Build the sample from categories instead: a few ordinary cases, the shortest and longest items, one with a missing field, one in an unusual format, one that a person would find genuinely ambiguous, and one where the correct answer is that the task does not apply. Twenty of those are worth two hundred convenient ones.
Keep that set. It becomes the fixed sample you re-run whenever the process changes, and it is the only way to tell whether a correction improved things generally or merely fixed the case you were annoyed by that morning.
Decide what counts as correct before you look
Grading a pilot after the fact invites you to accept whatever came back. Writing down beforehand what an acceptable result looks like — the fields that must be present, the tolerance on a judgement, what a failure would look like — converts a vague impression into a count.
Mark each result as correct, wrong, or needing a person, and count the third category separately. That number is usually the decisive one, because a process that produces good results for eighty per cent and requires human attention on the rest still requires a person to read all of it to find out which is which.
Where you can, have someone who did not build the process do the grading. The person who wrote the instruction reads the output charitably, having watched it improve, and their sense of whether it is working is not a reliable measurement.
The pilot answers the economics as well
A small run gives you real figures for the time and cost per item, which multiply into a number you can compare against doing the work another way. That comparison is frequently uncomfortable and it should be made explicitly rather than assumed at the start.
Add the cost of checking to the arithmetic, since it is part of the process rather than an afterthought. If verification takes as long as the work would have taken unaided, the honest conclusion is that this task is not a good candidate, however impressive the individual results look.
It also surfaces the practical obstacles that never appear at three items: rate limits, a source format that varies more than expected, items that time out, and the awkward question of what to do with the failures. Those become process design rather than surprises during a live run.
Stage the rollout even after the pilot passes
The gap between twenty and two thousand is large enough that a middle step is worth taking. A run of a couple of hundred surfaces the rare cases the small sample missed, and the rare cases are where the embarrassing failures live.
Make the first real run reversible where the system allows it. Write results to a staging location, produce a report rather than an update, or mark everything created by the process so it can be identified and removed as a group. That marking is a small piece of work that has saved a great many afternoons.
And keep looking after it starts working. The rate at which items need human attention is the metric worth watching, because it moves when the input material changes, and a process that quietly starts failing at a higher rate looks exactly like one that is still fine until somebody counts.
Common questions
How large should the pilot be?
Large enough to include the awkward cases rather than large enough to be statistically impressive — usually fifteen to thirty deliberately chosen items covering the extremes, the malformed, the ambiguous and the ones where the task does not apply. Composition matters far more than size.
What is the most important number to come out of a pilot?
The proportion of results that need a person to look at them. If that is high, the batch has not saved the work, because somebody must read everything to find out which results are in that group.
Should the person who built the process grade the results?
Preferably not, or at least not alone. Having watched the output improve, they read it charitably and grade against their expectations rather than against the requirement. A second reader with the written criteria produces a more useful count.
Features writer, Prompt After Prompt
Aarav covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.