Measure the whole task, or you will keep believing the fast part

The visible saving and the invisible cost
A draft that used to take ninety minutes now takes four, and that number is memorable because you watched it happen. The forty minutes spent afterwards checking figures, rewriting two sections and restoring the qualifications that vanished aren’t memorable, because editing has always been part of writing.
This asymmetry is why almost every informal estimate of time saved is too generous. The saving is concentrated, observed and surprising; the cost is diffuse, expected and spread across the rest of the day. Both are real, and only one gets counted.
The result is a process that feels transformative and delivers something more modest. That doesn’t make it worthless — a genuine third off a recurring task is worth having — but the difference between a third and the claimed nine tenths determines what else you should be building.
Measure from request to accepted result
The only honest boundary is the whole task, from the moment the work arrives to the moment somebody accepts the finished thing. Everything inside that boundary counts: the prompting, the waiting, the failed attempts, the checking, the corrections and the second round after a colleague sent it back.
Failed attempts are the most commonly excluded item and often the largest. Three regenerations before something usable is twelve minutes that nobody records, because each one felt like a moment rather than a period.
Rework belongs inside the boundary too, even when it happens days later. A piece returned by a reviewer for an error that came from the draft is part of the cost of that draft, and treating it as a separate event is how a process keeps its reputation despite its results.
Compare against a real baseline
Most comparisons are made against a remembered version of the old way, and memory of how long things used to take is unreliable in a predictable direction. It also tends to compare against doing the task well, when the honest alternative was often doing it briefly or not at all.
Where a decision matters, timing a handful of tasks the old way is worth the inconvenience. Where it doesn’t, at least be explicit about what the alternative actually was, because a great deal of generated output replaces something that would never have been produced, and that is a different kind of benefit from a saving.
Quality has to enter the comparison as well, and it is the part everyone skips because it is harder to count. A faster process that produces work a reviewer rates as slightly worse has not saved time in any meaningful sense; it has traded something for something, and the trade should be stated rather than hidden inside a time figure.
The clean cases are the ones where quality is fixed by a standard — the output either passes validation or it does not — which is one more reason to like tasks with checkable results.
Some tasks get slower and it is worth knowing which
A short piece you could have written in five minutes takes longer through a process that involves specifying, generating and reading. Tasks requiring context that only you hold take longer because supplying that context is most of the work. Anything where verification is expensive can easily cost more than it saves.
These are not failures of technique. They are the boundary of where the approach applies, and finding them is the main practical benefit of measuring anything at all. A team that knows which five tasks are slower has learned something more useful than a headline number.
The instinct to use the tool everywhere once it has worked somewhere is strong, and it is the specific habit that measurement corrects.
Keep the measurement cheap and occasional
None of this justifies a timekeeping regime. Timing a handful of representative tasks twice a year, honestly, tells you nearly everything a continuous system would, at a fraction of the cost and with a better chance of being done at all.
Write down the number alongside what was measured and what the alternative was, because the figure is meaningless without both. Six weeks later, when someone asks whether it is working, that note is the difference between an answer and an impression.
And treat the finding as provisional. The tools change, the tasks change, and a measurement taken a year ago describes a situation that may no longer exist in either direction.
Common questions
Why do informal time savings look so large?
Because the saving is concentrated and observed while the cost is diffuse and expected. The fast draft is memorable; the checking, the failed attempts and the later rework get absorbed into ordinary work and never counted against it.
What should be inside the measurement?
Everything from the arrival of the task to acceptance of the finished result — prompting, waiting, discarded attempts, checking, corrections and any rework triggered days later. Discarded attempts and delayed rework are the two most commonly excluded, and often the largest.
Which tasks tend to get slower?
Short ones you could do directly, ones needing context only you hold, and anything where verification is expensive. Identifying those is the main practical payoff of measuring, because the instinct after one success is to apply the approach everywhere.
Deputy editor, Prompt After Prompt
Adrian has written about prompt craft, writing with ai, images & audio for most of the last decade and prefers a plain explanation to a clever one.