Skip to content
Working with the tools, not about them

Images & Audio

Asking a tool to describe a picture is reliable about the subject and not about the detail

Reading images rather than generating them is a genuinely useful capability with a specific failure pattern, and the pattern determines which jobs it can be trusted with.

By Naina Sethi3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The general description is the trustworthy part

Supply a photograph and ask what it shows, and the account of the subject and the scene will usually be right. Two people at a table in a cafe, an industrial building photographed from the road at dusk, a spreadsheet on a laptop screen. That level of description is dependable enough to build on.

What comes with it, unrequested and in the same confident register, is detail. The colour of a jacket, the number of chairs, the words on a sign, the make of a machine, the mood of the person on the left. Some of that will be correct. Some of it is what usually accompanies scenes of this kind.

The two arrive in one paragraph with nothing to separate them, which is the difficulty in a sentence. The description is not uniformly reliable and does not indicate where the reliability drops.

Ask narrowly, and ask about presence rather than identity

Requests targeted at one property behave far better than a general description. Is there text in this image. How many people are visible. Is the room occupied. What is in the foreground. A narrow question limits the invented material because there is less room for it.

Presence questions are more reliable than identity questions. Whether a sign exists in the frame is easier than what the sign says; whether a person is present is easier than who they are or what they are feeling. Sort your questions by that distinction before deciding how much to trust the answers.

Where the answer must be exact — a serial number, a reading on a gauge, a handwritten note, a total on a receipt — treat the output as a first pass to be confirmed by a human looking at the image, particularly when the field is short and a single character changes the meaning.

Alt text is a good fit, with the substance supplied by you

Describing images for readers who cannot see them is one of the better uses, because the alternative in most organisations is nothing at all, and a serviceable description beats an empty attribute. Generated first drafts get the scene right and need editing for two things.

The first is length and purpose. Good alternative text says what the image contributes in its context rather than cataloguing what is in the frame, and only you know why the picture is on the page. A decorative image needs almost nothing; a chart needs its point stated, not its shape described.

The second is the invented detail, which is worse here than elsewhere, because the reader has no way to check it. A description containing a colour that is wrong or a sign that does not exist is a private error that the person relying on it cannot detect. Read every generated description against the picture before it ships.

Cataloguing and triage are where volume pays

Large image collections are the case where reading pictures earns its place. Rough tags, a one-line description, a shot type, whether a picture contains people: these are useful at a scale where nobody was going to write them by hand, and the errors are tolerable because the output is a search aid rather than a record.

Design it as triage rather than as description. The job is to make a collection navigable so a person can find candidates quickly, and it is worth stating that in the metadata itself so that nobody later mistakes generated tags for catalogued fact.

Sample the results as you would any batch process, and sample the ones that passed rather than only the ones that failed, since the failure here is a confident tag on the wrong picture rather than an obvious gap.

What it cannot tell you about the picture

It cannot tell you whether an image is authentic, whether it has been altered, where or when it was taken, or whether the scene it shows happened as it appears. Those are provenance questions, and they are answered by metadata, sourcing and the person who supplied the file, not by looking harder at the pixels.

It also should not be used to identify individuals, and a confident identification of a person from a photograph is both unreliable and a use with consequences that a description task does not carry.

For anything medical, forensic, safety-related or legal, a description of an image is not an assessment of it, and the distance between those two is exactly where somebody gets hurt. Use it to find the file. Have the file read by whoever is qualified to read it.

Common questions

Which parts of an image description can be trusted?

The subject and the scene are usually right. Fine detail — colours, counts, text on signs, makes and models, the emotional state of a person — is mixed in with the same confidence and includes material that typically accompanies such scenes rather than material that is present.

Why is invented detail worse in alternative text?

Because the reader relying on it cannot check it against the picture. A wrong colour or a sign that is not there becomes an undetectable error for the one audience that has no recourse, so every generated description should be read against the image first.

Can it tell you whether a photograph is genuine?

No. Authenticity, alteration, place and date are provenance questions answered by sourcing and metadata. Identification of individuals should not be attempted, and medical, forensic or safety assessments require somebody qualified to read the image.

Images & Audioimage readingalt textverificationaccessibility
Naina Sethi
Consumer editor, Prompt After Prompt

Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.