An image prompt describes a picture, not the scene it depicts

Narrative and depiction are different specifications
Ask for a woman who has just been told bad news and you will get a woman looking sad, more or less generically, in a setting the tool chose. Nothing in the request said where the camera is, how much of her is in frame, what the light is doing, what she is standing in front of, or whether we can see her face at all.
A photographer given the same brief would answer all of those before pressing anything, and the answers are the picture. The emotional content of an image comes from framing and light far more than from an expression, which is why the generic sad face is the least effective way to convey the thing you asked for.
So the first correction, and the one that changes results most, is to stop describing the situation and start describing the frame. What is in it, from where, at what distance, lit how.
The four decisions that carry most of the result
Subject and setting are obvious and usually the only ones people supply. The other three are shot size, viewpoint and light. Shot size — close, waist-up, full figure, wide — determines what the image is about, since anything not in frame is not in the picture regardless of how carefully you described it.
Viewpoint decides the relationship. Slightly below the subject makes them dominant, slightly above makes them vulnerable, and eye level makes them ordinary. These are not artistic flourishes but the basic grammar of pictures, and they are cheap to specify in three or four words.
Light is the one that separates a flat result from a convincing one. Hard light from one side, overcast and even, backlit with the subject in shadow, a single lamp in a dark room: each of these produces a different image from identical content, and stating one prevents the tool from defaulting to the bright, even, characterless lighting that most generated pictures share.
Name the medium, because it constrains everything else
Whether the result should look like a photograph, a pencil drawing, a woodcut, a flat vector illustration or a painting is the largest single decision, and leaving it out means inheriting whatever the rest of the wording implies. A prompt full of photographic words produces something photographic by accident rather than by choice.
Medium also carries a set of physical consequences that you then do not have to state. Ask for a photograph with a shallow depth of field and the background handles itself. Ask for a linocut and you get the limited palette, the cut edges and the flat areas without listing them.
This is a general principle worth extracting. One well-chosen constraint that implies a dozen others is more effective and more stable than a dozen separate adjectives, which tend to compete with each other and dilute.
Composition instructions work better as facts than as effects
Asking for a balanced composition or a dynamic angle asks for a judgement about the finished image, which is not something a description can specify directly. Asking for the subject placed to the left of frame with the doorway behind them on the right is a fact about the picture, and facts get rendered.
The same applies to mood. Atmospheric, moody and cinematic are among the most-used words in image prompting and among the least precise. Each of them is produced by concrete choices — low light, deep shadows, a restricted palette, haze in the air — and naming those gets you a specific version rather than the average of everyone else’s.
Where relationships between elements matter, keep them simple and few. Precise spatial arrangement of several objects remains one of the weaker areas across these tools, and a prompt describing five things in specific positions relative to each other will usually get two of them right.
Know when the description has stopped being the tool
There is a point in any image task past which words are the wrong instrument. If you need this specific person, this exact product, this arrangement, or a change confined to one corner of an otherwise finished picture, describing it again more carefully will not converge. Editing tools, reference images and masking exist because description does not reach that far.
The honest boundary is roughly this: description is excellent at establishing a look and poor at achieving a target. If you can accept a range of acceptable outcomes, generation is fast and often better than what you had in mind. If only one outcome will do, budget for the tool being the wrong choice.
A useful discipline is to decide which category you are in before you start, because the frustrating sessions are almost always the ones where a target was needed and a range was on offer.
Common questions
Do longer image prompts work better?
Only up to a point, and past it they get less predictable rather than more. A long list of adjectives dilutes the influence of each, and terms that pull in different directions produce averaged, muddled results. A compact prompt covering medium, subject, framing, viewpoint and light usually outperforms a long one.
Why does the same prompt give such different results between tools?
Because each tool was fitted to a different collection of images with different captions, so the same words map to different regions. Style vocabulary is the least portable part; concrete descriptions of framing, light and medium travel considerably better, which is a reason to prefer them.
Is there any point specifying camera settings?
They act as shorthand for a look rather than as instructions to a camera. Naming a long lens or a wide aperture reliably shifts compression and background blur, because those associations are strong in captioned photographs. Naming an exact shutter speed generally does nothing, since it has no consistent visual signature.
Consumer editor, Prompt After Prompt
Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.