Skip to content
Working with the tools, not about them

Images & Audio

Starting from a photograph you already have is a more controllable problem

Supplying an existing picture as the basis of a generation replaces the hardest part of image prompting, which is describing a composition precisely enough to obtain it.

By Naina Sethi3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Description is the bottleneck, and an image removes it

Getting a specific arrangement out of a written prompt is slow. Three objects in a particular relationship, a subject at a particular angle, a view from a particular height — each has to be described in language, and language is a lossy way to specify geometry. Most of the iteration in image work is spent here.

A supplied picture states all of it at once. The angle, the proportions, the arrangement and the amount of empty space are given rather than described, and the request that remains is about treatment: change the season, restyle the surface, replace the background, render it differently.

This is why a bad phone photograph is often a better starting point than a paragraph of careful description. It does not need to be attractive. It needs to be correct about the structure you want preserved, because that is the part the written prompt was failing to convey.

How much of the original survives is the main control

Every tool of this kind exposes some way of setting how closely the result follows the input, whatever it is called. Set it high and you get a lightly modified version of your picture; set it low and the input becomes a loose suggestion, at which point you are back to prompting with extra steps.

The useful working method is to find the setting where structure is retained and surface is free. That point differs by tool and by image, and locating it takes a handful of test generations, which is a much better use of attempts than rewriting a description repeatedly.

It is worth knowing what the input does not fix. Colour and lighting often shift more than expected even at high fidelity, and small details are regenerated rather than preserved. If a particular detail must survive exactly, it belongs in a masked edit or an ordinary editor rather than in the generation.

The permission question comes first

The input is somebody’s picture unless it is yours. Using a photograph you found as the basis for a generated image is a different act from being inspired by it, and depending on where you are and what you are making, it may not be something you are entitled to do.

People are the sharper case. A recognisable face carries interests that have nothing to do with copyright, and generating variations of a person from a photograph — a colleague, a customer, anyone who did not agree to it — is the kind of thing that damages trust badly and is regulated differently in different places.

The safe practice is boring and effective. Use your own photographs, pictures your organisation owns, or material licensed for the purpose, and get explicit agreement before using anyone’s likeness. This is not a technical constraint and no setting in the tool will resolve it for you.

Rough input, precise output

The technique extends past photographs. A crude sketch, a grey-box arrangement of shapes, a screenshot of a layout, a scribbled diagram of where things should sit — any of these communicates structure faster than a paragraph, and none of them requires drawing ability.

This is where the approach earns its place in ordinary work. Blocking out a composition in five minutes with rectangles, then generating a finished treatment over it, gives you compositional control that no amount of describing achieves, and it keeps the design decisions with the person making them.

Combining an input with the region-limited editing available in most of these tools covers most practical needs: supply the structure, generate the treatment, then repair the two or three places that came out wrong. Each step is bounded, which is what makes the whole process predictable.

When the photograph should just be the photograph

If you already have a usable picture of the actual thing, generating a version of it is often a loss. Real products, real premises, real people and real events carry credibility that a synthesised version does not, and readers are increasingly quick to notice the difference.

There is also a category where the transformation misleads. Making a room look brighter than it is, a product cleaner than it arrives, a venue larger than it is — these are commercial claims made in pictures, and the fact that a tool made them easy does not make them acceptable.

The honest use is closer to illustration than to documentation. Where the picture stands for an idea, transforming a reference gives you control and speed. Where the picture is evidence about something real, the camera remains the correct instrument and no prompting technique changes that.

Common questions

Does the reference picture need to be good?

No. It needs to be structurally correct — right angle, right arrangement, right proportions — because structure is what it contributes. A rough phone photograph or a sketch of grey boxes conveys composition far more precisely than a paragraph of description.

How do I keep a specific detail from being changed?

Raise the fidelity to the input, and if it still shifts, take that detail out of the generation entirely. Masked editing or a conventional image editor preserves exact pixels, whereas any generation regenerates small detail even when the overall structure is held.

Is it acceptable to use a photograph I found online as an input?

Often not. Using someone’s picture as the basis for a generated image is a use of that picture, and a recognisable person brings further interests beyond copyright. Rules vary by country and by use, so work from material you own or have licensed, and get agreement before using anyone’s likeness.

Images & Audioreference imagescompositionphotographycontrol
Naina Sethi
Consumer editor, Prompt After Prompt

Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.