Anything that has to be counted should not be asked of an image generator

The failures cluster, and the cluster is predictable
Exact quantities. Legible text of any length. Precise spatial arrangements. The same character appearing twice with the same face. Hands, though less reliably than the joke suggests. Anything where a small local detail must be exactly right rather than plausibly right.
These are not a random collection of bugs waiting to be fixed one by one. They share a common shape: each requires the image to satisfy a discrete, checkable condition, and the process producing the image is optimising for overall plausibility rather than checking conditions. Something that looks like five apples is a success by that measure.
Improvement across generations of these tools has been real, particularly on text rendering, and the class of failure remains. Treat specific claims about what is now fixed with mild scepticism and test rather than assume.
Design the requirement out of the picture
The reliable move is to remove the discrete condition from the generated part of the job. If the image needs a headline on it, generate a picture with clear space and set the type yourself in a layout tool, where it will be correct, editable and in the right font. This is faster than the prompting alternative and it always works.
If a specific count matters, crop to it or compose it. Three items in frame is achievable by generating a larger group and cropping, or by generating each and assembling. If a diagram has to be accurate, it is not an image generation task at all; diagrams are made in diagram tools.
This is the general principle for working with any tool that is strong on plausibility and weak on precision. Give it the part where a range of outcomes is acceptable, and keep the part where only one outcome is correct.
Consistency across images is its own problem
Wanting the same character, product or setting across a series is one of the most common practical requirements and one of the least well served by description alone. Two prompts with identical wording produce two different people, because nothing in the wording pins down a face, and a face is a great many small decisions.
The mechanisms that exist for this — reference images, subject conditioning, whatever a given tool calls it — work considerably better than words and still drift, particularly across changes of pose, expression or lighting. Plan for a tolerance rather than for identity, or design the series so that the character is seen from behind, at distance, or partially.
For anything commercial with a recognisable person or product involved, this constraint often decides the approach entirely, and it is better to discover that at the planning stage than after a day of generation.
Checking is the step people skip
Generated images invite a glance rather than an inspection, because they read as finished. The errors that survive to publication are almost always in the parts nobody looked at: a background sign with plausible non-words, a reflection that does not match, a hand with an extra joint, a chair leg that terminates in nothing.
A deliberate pass helps, and it is worth doing in a fixed order so that it does not depend on attention. Edges of frame. Background text and signage. Hands, feet and where limbs meet the body. Reflections and shadows against the light direction you asked for. Repeated elements, which is where duplication artefacts hide.
View at full size for this. A small preview hides exactly the class of error that is most embarrassing at scale, and the image will be seen at whatever size your reader chooses rather than the one you checked at.
It also helps to look at the picture the way a stranger will rather than the way you made it. Flipping it horizontally makes errors of symmetry and anatomy jump out, because the familiarity that was hiding them is broken. Coming back an hour later does something similar and costs nothing but patience.
When to conclude this is the wrong tool
The decision is easier than the endless-refinement habit suggests. If the brief contains a hard constraint — this exact logo, these exact words, this specific person, an accurate technical illustration — the generator is not the instrument, and the time spent trying to make it one is time that would have finished the job another way.
Where the brief is a mood, a texture, a background, a starting point for an illustrator, or one of forty images nobody will study, it is very good and extremely fast. The competence is real; it is just narrower than the demonstrations imply.
The people who get consistent value from these tools are mostly the ones who decided early which half of that split a job falls into. It is the least technical skill in the whole area and probably the most useful.
Common questions
Has text rendering been solved?
It has improved a great deal and short phrases are now often correct in the better tools, which was not true a couple of years ago. Longer text, specific typefaces and small text remain unreliable. For anything that has to be exactly right, adding the type in a layout step is still the sensible route.
Why are hands specifically such a problem?
They combine several of the hard properties at once: a variable number of parts that must be counted, articulation in many directions, frequent self-occlusion, and enormous variation in the source material. The result is a region where plausible-looking output is easy and correct output is hard.
Can inpainting fix these errors?
Often yes, and it is the right first response to a local defect in an otherwise good image. Mask the region, regenerate just that part, repeat if needed. It fails where the error is structural rather than local, and where fixing one area disturbs the surrounding context enough to create a new problem.
Consumer editor, Prompt After Prompt
Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.