Skip to content
Working with the tools, not about them

Workflows

One long instruction hides which step failed

A single prompt that extracts, judges and writes gives you no way to tell which of those three went wrong, and the fix is to make each one produce something you can look at.

By Bhavna Deshpande4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A compound task fails as a black box

The instruction that reads a support thread, decides whether the complaint is justified, and drafts a reply is doing three unrelated jobs. When the reply is wrong, you cannot tell whether it misread the thread, applied the wrong standard, or understood everything and wrote it badly. All you have is the final text.

So the response is to adjust the wording somewhere and try again, which is guessing. Sometimes the guess works and you have learned nothing transferable; more often you cycle through several rewrites while the actual fault sits in a part of the task you never examined.

Splitting the job is not about making each piece easier for the tool, though it usually does that too. It is about making the failure visible, which is the difference between a process you can improve and one you can only re-roll.

Intermediate outputs are the point

When the extraction step produces a list of the claims made in the thread, you can read that list. If it is wrong, you know exactly what to fix and you know that everything downstream was working from bad material. If it is right, the fault is later, and you have halved the search without any cleverness.

This has a second benefit that matters more over time. An intermediate artefact can be checked mechanically. A list of extracted dates can be validated against a format; a set of quoted lines can be confirmed to appear in the source. Neither check is possible on a finished paragraph that has already absorbed and reworded everything.

Design the intermediate steps to produce structured output for exactly this reason. The shape you would choose for readability and the shape you would choose for checking are usually the same shape.

Chains have their own failure mode

Splitting a task is not free. Every step introduces a place where a small error becomes an input to the next one, and errors compound rather than cancelling. A five-step chain where each step is nearly always right can still be unreliable end to end, and the arithmetic is unforgiving in a way that intuition is not.

The mitigation is to keep chains short and to put the fragile steps early, where an error is still cheap to catch. Anything involving judgement or inference should be as close to the source material as possible, since each subsequent step is working from a summary of a summary.

It is also worth asking which steps need a tool at all. Plenty of chains have a middle step that is really a filter, a lookup or a format conversion, and those are more reliably done with ordinary code that fails loudly rather than plausibly.

Where a single prompt is the right answer

Decomposition is not automatically better. For a small, well-defined task with an output you can verify at a glance, one instruction is simpler, faster and has fewer places to go wrong. Splitting it adds machinery to maintain for no gain in reliability.

The threshold is roughly this: split when you cannot tell from the output which part went wrong, or when a step needs checking before the next one runs, or when the same intermediate result would be useful for more than one downstream job. Below that, leave it alone.

There is also a real cost to over-engineering. A five-stage pipeline for something done twice a month is a maintenance burden that will quietly rot, and the person who inherits it will not know which stages still matter.

A reasonable compromise for medium-sized tasks is to keep one instruction while requiring the intermediate result to appear as a separate, labelled part of the output. You get something to inspect without building a pipeline, and if one stage turns out to be the persistent problem, you already know which part to split off properly.

Naming what each step is responsible for

The habit that makes any of this work is writing down, in one sentence, what a step is for and what it must never do. Extract, do not interpret. Judge, do not write. Draft from the supplied judgement, do not revisit it. Steps that stay inside their remit are the ones that stay debuggable.

It also prevents the most common drift in maintained workflows, where a step gradually absorbs responsibilities from its neighbours because it was easier to add a clause than to change the structure. Six months later there are three steps and one of them does everything.

The parallel with ordinary software is not a coincidence, and neither is it exact. Deterministic components can be tested once; these cannot, since the same input may not give the same output. That difference is why the visible intermediate result matters so much more here than it does in code you have written yourself.

Common questions

Should each step be a separate conversation or one long one?

Separate, wherever the tooling allows it. A single conversation carries everything said earlier into every later step, which is exactly the coupling you were trying to remove, and it makes an individual step impossible to test in isolation. Pass the intermediate output forward explicitly instead.

Does splitting cost more?

Usually a little more in total processing, since context gets re-supplied to each step, and considerably less in your own time once anything goes wrong. For work that runs at volume the arithmetic is worth checking. For work done by a person at a keyboard, the time saved debugging dominates easily.

How do I know how many steps to use?

Start with one and split at the point where you cannot diagnose a failure. Adding steps speculatively produces pipelines with stages nobody can justify. Every split should be traceable to a specific occasion when you could not tell what had gone wrong.

Workflowsworkflowchainingdebuggingsteps
Bhavna Deshpande
Editor, Prompt After Prompt

Bhavna covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.