A process that runs unattended decays without announcing it

Nothing in the process notices that the world moved
A workflow is built against a particular set of circumstances: a source document with a certain layout, a tool that behaves a certain way, a set of categories that covered the material at the time. None of those are pinned. All of them change, usually without an announcement, and the process carries on regardless.
The failure mode is not a crash. A crash would be a gift, because somebody would investigate. What actually happens is that the output stays well formed and becomes wrong — extractions that pick up the header instead of the value, classifications that put everything into one category, summaries that quietly stop covering the last section.
And because the output looks the same as it did last month, nobody rereads it. The people downstream trust it precisely because it has been reliable, which is why the discovery is usually made by an outsider, several weeks later, in an awkward setting.
The assumptions that expire are predictable
Input format is first. A supplier changes a report layout, a system adds a column, an export starts including a footer, and a step that located a value by position begins locating something else. Anything that depends on the shape of an incoming file is exposed.
The tool underneath is second. Models are updated, defaults change, a response becomes more verbose or a formatting habit shifts, and an instruction tuned around the old behaviour stops producing the old result. Nothing on your side changed, which is what makes this one so confusing to diagnose.
The third is your own categories. A classification scheme written a year ago reflected the material then, and new kinds of item now arrive that do not fit any of the options. They still get classified, into whichever bucket is closest, and the resulting figures are wrong in a way that is invisible from the totals.
Check the assumptions, not only the output
The cheapest safeguards are mechanical and specific. Does the input have the expected number of columns and the expected header text. Is the output length within the usual range. Did any required field come back empty. Did more than the usual proportion fall into the catch-all category. Each of these is a few lines and each catches a different silent failure.
Distribution checks are underrated. If eight per cent of items normally land in a particular category and this week it is forty, something has changed even though every individual result looks reasonable. Aggregate patterns reveal drift that item-level inspection does not.
Sampling remains necessary alongside the automatic checks. Reading five results properly every week, chosen at random rather than from the top, catches the class of problem that no rule anticipated, and it takes ten minutes.
Fail loudly and stop
The default behaviour of a lot of homemade automation is to carry on when something is missing, because that seemed robust at the time. It is the opposite of robust. A process that skips malformed items silently produces a short output nobody counts, and one that substitutes a default value produces a plausible wrong figure.
A better default is to stop and say why, with the offending item attached. Interruption is annoying and it is much cheaper than a month of quietly wrong records, particularly for anything that writes into a system other people read.
Keep enough of a record that a problem can be reconstructed: what ran, when, against what input, with which version of the instruction, and what came out. Without that, diagnosing a drift means guessing about a state that no longer exists anywhere.
Schedule the review, because nobody volunteers for it
Working automation attracts no attention, so its review has to be arranged in advance. A recurring calendar entry to re-run the saved sample against the current process and compare against the stored results is the whole practice, and it takes half an hour.
The same review is the moment to ask whether the process is still needed, and whether the reason for each of its steps still applies. Workflows accumulate stages that were added to work around a limitation that no longer exists, and nobody removes them because nobody is quite sure what they do.
It is also worth naming an owner. An automated process with no owner is one that will be maintained by whoever is nearest when it breaks, which is a poor arrangement for something producing figures that other people rely on. If nobody will own it, that is a reasonable argument for not running it unattended at all.
Common questions
Why does a workflow break when I have not changed anything?
Because the things it depends on are not under your control. Input layouts change, tools are updated and their default behaviour shifts, and the categories you defined stop covering the material. The instruction is the same; the conditions it assumed are not.
What is the cheapest useful monitoring to add?
Assumption checks on the input — expected columns, expected headers — plus range checks on the output and a count of how many items land in the catch-all category. Watching that distribution over time catches drift that reading individual results does not.
Should a process stop when it hits something unexpected?
For anything writing into a system other people read, yes. Skipping malformed items quietly produces a short output nobody counts, and substituting defaults produces plausible wrong figures. An interruption is cheaper than weeks of undetected errors.
Editor, Prompt After Prompt
Bhavna covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.