Skip to content
Working with the tools, not about them

Prompt Craft

An instruction that only works once was underspecified, not unlucky

The gap between a prompt that produced something good and a prompt you can rely on is a list of decisions you left open without noticing.

By Adrian Novak4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The result you liked was one draw from a range

The usual sequence is this. You type something quickly, the answer comes back better than you expected, you save the wording somewhere and treat it as a tool. A week later the same wording on a similar job produces something flat, oddly formatted and half the length, and the obvious conclusion is that the tool has got worse.

It generally has not. The first answer was one sample from the range of things your instruction permitted, and it landed at the good end. Nothing about that instruction made the good end more likely than the mediocre one; you simply saw the good end first and took it as the baseline.

Two sources of variation are in play, and they are not the same size. There is the sampling inside the tool, which you can sometimes reduce and never eliminate. And there is every decision your wording left open, which gets made freshly on each run. The second is far larger, and it is entirely yours to fix.

Read your own prompt as if you had to carry it out

Hand the wording to someone who does not know the job and listen to the questions. How long should it be? Who is reading it? What if the document does not actually contain the answer? Should the figures be rounded? Do you want headings, and in what order? Each question marks a decision the prompt delegated silently.

Prompts that behave consistently tend to pin down four things: who the output is for, how long it should be, what structure it takes, and what to do in the awkward case. The last one is skipped most often and causes the worst variance, because an unhandled edge is exactly where output stops being merely wrong and starts being invented.

There is a related failure that looks like precision. Three careful sentences about tone and nothing at all about structure will give you beautifully worded output with a different shape every time. Specificity in the wrong place feels like effort and buys nothing.

Test it the way you would test anything else

A single run tells you almost nothing. Run the same instruction three times against the same input and you learn how much room it left. Run it once each against three genuinely different inputs and you learn whether it generalises. Those are separate questions, both cheap to answer, and most people answer neither.

Keep the difficult inputs somewhere. A source that is unusually short, one in a format you did not anticipate, one where the answer is genuinely absent, one twice the normal length. An instruction that handles the median case is a demonstration rather than a component you can build on.

You are not looking for perfection. You are looking for a failure you can predict. Something that goes wrong the same way each time can be corrected with a sentence. Something that goes wrong differently each time has not been specified yet, and adding emphasis will not help.

Some of the variation is not yours to remove

Even a tightly written instruction wanders on genuinely open tasks. Ask for a summary of a long report and the emphasis will shift between runs, because choosing what matters is a judgement and no wording turns a judgement into a lookup. Chasing determinism there is wasted effort and produces stilted, over-constrained output.

The workable split is between variation that costs you something and variation that does not. Length, structure, what gets included and what gets left out: pin those, because a downstream step or a reader depends on them. Which of two decent phrasings appears: let it go, and stop reading the difference as a signal.

Underneath all of this the tools themselves change. An instruction that depends on a particular quirk of phrasing — the elaborate threats, the strange incantations people trade — tends to stop working first. Plainly stated requirements survive version changes considerably better, which is an argument for boring prompts.

A durable instruction looks unremarkable

What you end up with is dull to read: the task, the audience, the constraints that matter, a rule for the awkward case, and the shape of the output. Fewer adjectives than you would expect. More nouns. Very little of the performative framing that circulates as advice.

It also lives in a file rather than in chat history. Wording you cannot retrieve is not a process, it is a memory of one, and the first time you need to change it you will rewrite it from scratch and lose whatever you had learned.

The shift that makes the difference is small but hard to unlearn. You are not asking for something and hoping. You are writing a specification that a competent but literal-minded stranger will follow, once, without being able to ask you anything.

Common questions

Does turning the randomness down make prompts reliable?

It reduces one source of variation and leaves the bigger one untouched. Lower randomness means the tool picks the most likely continuation more consistently, but if your instruction left the length or structure open, the most likely continuation still depends heavily on the input. It is worth doing for extraction and formatting work, and it will not rescue a vague brief.

How many test inputs are enough before I trust a prompt?

For anything that runs more than a handful of times, three to five inputs chosen to be different from each other will surface most problems, and one of them should be a case you expect to fail. This is not statistics, it is sanity checking. The point is to see the failure shape before it appears in something you have already sent.

Should I write long prompts or short ones?

Long enough to close the decisions that matter and no longer. Padding an instruction with restatements and emphasis tends to dilute the parts that are doing work. If you cannot say what a sentence in your prompt is preventing, it is probably not preventing anything.

Prompt Craftpromptingrepeatabilityspecificationtesting
Adrian Novak
Deputy editor, Prompt After Prompt

Adrian has written about prompt craft, writing with ai, images & audio for most of the last decade and prefers a plain explanation to a clever one.

Prompt Craft

Give it the brief, not the archive

Supplying background is a selection problem, and the instinct to include everything relevant produces worse…

· 3 min read