Skip to content
Working with the tools, not about them

Workflows

The step worth automating is the one whose failure you would notice

Automation decisions here turn less on how well the tool performs a task than on whether a wrong result would be visible before it mattered.

By Aarav Sinha4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Capability is the wrong first question

The usual approach is to ask whether a tool can do a task, which is answered by trying it a few times. The answer is often yes, and it is not the question that determines whether automating it is a good idea. What matters is what happens on the occasions when it does the task badly.

Two steps with identical success rates can be completely different propositions. Drafting an internal summary that a knowledgeable colleague will read is safe, because a wrong summary is obvious to its reader and costs a minute. Categorising records that feed a report nobody re-reads is not, because a wrong category is invisible and permanent.

So the question to ask about each step is how a failure would surface, how long it would take, and what it would have touched by then.

The properties that make a step a good candidate

Frequency is the obvious one, since anything done twice a year is not worth the machinery. Beyond that, the useful properties are that the output is checkable by rule, that the input is reasonably uniform, that an error is recoverable, and that the step does not require knowledge the tool has no access to.

That last one rules out more than people expect. A step needing the current state of a negotiation, an unwritten policy, or the reason a particular customer is treated differently will produce a confident and reasonable answer built on none of that. It looks like a capability problem and it is an information problem.

Uniformity of input matters more than it sounds. A step that handles the standard case well and produces confident output on the unusual one is a specific hazard, since the unusual cases are exactly the ones where the consequences of being wrong tend to be larger.

Reversibility deserves separate weight from cost. A step producing a draft that somebody will edit is recoverable at every point along the way. A step that sends a message, updates a record or moves money is not, and the reliability that felt perfectly acceptable in the first case should not be assumed acceptable in the second.

The trap of nearly finishing

There is a familiar shape where a process handles most cases well, the remainder need human attention, and the human cannot tell which is which without examining everything. The automation has removed the work of doing the task and added the work of checking it, and depending on the task those can be similar in size.

The way out is usually to make the process declare its own uncertainty in some usable form — flagging the cases it could not handle cleanly, or being designed so that the difficult class is routed elsewhere before it is attempted. Sorting the input is often more valuable than improving the handling.

Where neither is possible, the honest conclusion is that this task is not a good fit yet. That is a legitimate outcome, and treating it as one saves more time than another round of prompt refinement.

Keep the judgement, hand over the preparation

The division that holds up across most working contexts is that the tool prepares and a person decides. Gather the relevant material, extract the fields, produce the comparison, draft the options, flag the anomalies. Then a person who knows the context makes the call that has consequences.

This is not a moral position about human oversight, though there are those. It is that preparation is checkable and deciding often is not, and that the parts requiring knowledge outside the material are precisely the parts that cannot be supplied. Automating decisions while leaving preparation manual is the arrangement that reliably disappoints.

It also degrades gracefully. When a preparation step goes wrong, the person deciding usually notices that the material is odd. When a decision step goes wrong, the preparation was fine and nothing looks unusual at all.

Automate loudly

Whatever you do automate should fail in a way somebody sees. A step that quietly produces something plausible when its input was missing is worse than one that stops, because the failure enters the record and gets treated as data. Explicit refusal to proceed on unexpected input is a feature worth building deliberately.

Counting things helps as well. How many items ran, how many were flagged, how many a person overrode. A workflow with no numbers attached to it is one where a gradual degradation — a changed input format, a tool update, a source that stopped being maintained — will be discovered by a complaint rather than by you.

The general principle is old and it applies here with more force than usual, because the characteristic failure of these systems is producing something reasonable-looking rather than producing nothing. Silence is not evidence that anything is working.

Common questions

Is it worth automating something I only do occasionally?

Usually not as a pipeline, and often yes as a saved instruction. The documentation and the tested wording are where most of the value sits for infrequent tasks, and they cost almost nothing to keep. Building and maintaining automation for a rare task tends to result in machinery that has quietly broken by the next time you need it.

How do I decide what to check when a process runs at volume?

Check what a rule can check on everything, and sample the rest with attention to the unusual cases rather than to a random selection. Random sampling estimates an overall rate; deliberately looking at the odd inputs finds the failures that matter, since those are where confident wrong output concentrates.

Should a person always be in the loop?

It depends entirely on what an error costs and whether anyone downstream could detect one. For low-stakes, reversible, self-evident output, review is overhead. For anything entering a record, reaching a customer, or informing a decision the recipient cannot check, a person in the loop is the whole safeguard and removing it should be a deliberate, documented choice.

Workflowsautomationjudgementriskworkflow
Aarav Sinha
Features writer, Prompt After Prompt

Aarav covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.