Waiting time and cost per item are design decisions, not details

Per-item figures decide what is possible
A step that takes eight seconds is imperceptible while you are testing and becomes two and a half hours across a thousand items. A cost that is negligible per request becomes a line item somebody questions once it runs daily. Neither of these is a quality problem and both determine whether the process survives.
The arithmetic should be done at design time rather than discovered in the first month. Multiply the per-item time and cost by the real volume, add the checking, and compare it against the alternative, including the alternative of not doing the task at all.
It matters most where a person is waiting. An interactive step that takes twenty seconds will be abandoned, and a process nobody uses has an effective cost of whatever it took to build. Where the result is delivered to a person in real time, latency is a functional requirement rather than a preference.
Do the expensive thing once
A great deal of repeated work is genuinely repeated. The same document summarised for five different questions, the same reference material re-supplied on every request, the same classification recomputed on items that have not changed since yesterday. Each of these is an opportunity to do the work once and store the result.
Storing intermediate output is the simplest version and often the most valuable. Extract the structured facts from a document once, keep them, and answer subsequent questions from the extraction rather than from the original. The second question then costs almost nothing and returns immediately.
The discipline is knowing when a stored result has gone stale. Anything derived from a source that changes needs either a marker recording which version it came from, or a rule about how old a result may be before it is recomputed. A cache with no expiry policy eventually becomes a source of confident, out-of-date answers.
Match the effort to the difficulty of the item
Treating every item identically means paying the price of the hardest one across the whole set. Most batches are heavily skewed: a large majority of items are straightforward and a minority genuinely need care, and a design that recognises the difference is much cheaper than one that does not.
The usual arrangement is a cheap first pass that handles the clear cases and flags the rest. A simple rule, a keyword match or a fast lightweight step can dispose of a large share of the work, leaving the expensive, careful treatment for the items where it changes the answer.
This only works if the cheap pass is honest about its uncertainty. A first stage that classifies everything confidently defeats the purpose, so it must be allowed to say that it does not know, and the proportion it hands on is the number to watch when the input material changes.
Batching, parallelism and the limits on both
Grouping items into a single request reduces overhead and often cost, at the price of a longer response that has more room to drift in format and a single failure that takes the whole group with it. Groups of five to twenty are usually a reasonable compromise; hundreds are not.
Running many requests at once helps with total elapsed time and does nothing for cost per item. It also runs into rate limits, which is where a design that looked fine at ten items becomes unpredictable at a thousand, so any process intended to scale needs a retry policy that waits rather than hammers.
Both techniques add complexity that has to be maintained by somebody. For a job that runs monthly across two hundred items, sequential and simple is very often the better engineering, and the time saved by parallelising it would not repay the time spent building it.
Say when the arithmetic does not work
Some tasks are simply not worth automating at the volume they occur. Forty items a month, each needing careful review, is an afternoon of ordinary work and a fortnight of process design that then requires maintenance forever. The honest recommendation there is to do the work.
Others fail on the checking cost rather than the production cost. If verifying a result takes nearly as long as producing it unaided, the process has moved effort rather than saved it, and that comparison should be made with real figures from a pilot rather than from optimism.
The general shape of the answer is that these tools are cheap per item and not free, and that the cost of the surrounding process is usually larger than the cost of the requests. Designing as though the requests were the expensive part leads to elaborate optimisations around the wrong number.
Common questions
How do I decide between one large request and several small ones?
Grouping reduces overhead and increases the chance of format drift, and a failure takes the whole group. Small groups of five to twenty are a reasonable middle for most batch work, while anything a person is waiting for should be sized for response time instead.
Is caching results worth the complexity?
Where the same expensive step is repeated over unchanged input, yes, and it is often the single largest saving available. The requirement is a policy for staleness — a version marker or a maximum age — because a cache with no expiry produces confident answers from an outdated source.
When is automating a task not worth it?
When the volume is low enough that doing the work directly costs less than building and maintaining the process, and when checking the output takes nearly as long as producing it unaided. Both are answerable with figures from a small pilot rather than by argument.
Consumer editor, Prompt After Prompt
Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.