Skip to content
Working with the tools, not about them

Images & Audio

A generated clip is a shot, and a sequence of them is a different problem

Short moving images are produced one shot at a time, which means the difficulties are continuity, timing and revision cost rather than the description of any single frame.

By Devika Menon3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Length is the constraint everything else follows from

Generated motion arrives in short pieces, and whatever the current limit happens to be, the practical consequence is stable: you are producing shots, not sequences. Anything longer is assembled in an editor from separate generations, exactly as film has always been assembled.

That reframes the skill. The interesting decisions are which shots the piece needs, how they cut together, and what has to remain constant across them, none of which is a prompting question.

It also explains why an attempt to generate a whole scene in one go disappoints so consistently. A single continuous take of anything complicated is difficult to direct even with a camera and a crew, and asking for one from a description is a request for the hardest version of the job.

Continuity is the expensive part

Two clips of the same subject will differ in ways that are invisible in isolation and glaring in a cut: the shirt changes shade, the room is a slightly different room, the light comes from the other side, the person is a near relation of the person in the previous shot.

The mitigations are the ones an art department would recognise. Fix as much as possible outside the generated material: one location described identically every time, a written description of the subject reused verbatim, a consistent time of day and light direction stated in every prompt.

Where a tool allows a still image or a previous frame to seed the next clip, that will hold continuity better than words, and it will still drift over a sequence. Plan cuts that hide the drift — a change of angle, a cutaway, a shot of something else entirely — rather than expecting a run of matched shots.

The most reliable structural answer is to need fewer matched shots. A sequence built from distinct subjects, or from details rather than wide views, has almost no continuity burden at all.

Motion has to be described as motion

A prompt written as a still description produces something that barely moves, or moves in a slow drift that looks like a photograph being pushed around. Motion needs its own specification: what is moving, in which direction, how fast, and whether the camera is moving or the subject is.

Camera language works better than adjectives here because it is precise about what changes. A slow push in, a static frame with a subject crossing it, a handheld follow: each of these describes a relationship between the frame and the subject rather than a mood.

Keep one motion per clip. Compound instructions — the camera pans while the subject turns and the door opens behind them — are the video equivalent of asking for an exact count, and they fail in the same way and for the same reason.

Revision is regeneration, and it costs more than you think

A generated clip cannot be adjusted the way a still can be locally repaired. Changing one element usually means generating again, and generating again changes everything, so the small fix that would take a minute in a still image is a fresh draw of the whole shot.

This makes planning worth more here than in still work. Decide the shot list, the aspect ratio and the treatment before generating anything, because a late change to any of those invalidates the material rather than modifying it.

Keep every usable take rather than only the best one. Editing is where a piece is actually made, and an editor with four imperfect takes has options that an editor with one good one does not.

Sound is a separate job, and so is the question of what it depicts

Generated clips typically arrive silent or with sound that does not survive scrutiny, and the audio is built separately in the edit. Nothing about that is unusual; most film sound is constructed rather than recorded with the picture. It does mean the work is not finished when the pictures are.

The harder question is what the sequence appears to show. Moving images carry more evidentiary weight than stills, and a short clip of an event that did not happen is a considerably stronger claim than a photograph of the same thing. Anything depicting real people, real places or real events needs to be labelled plainly.

And for documentary, news or evidential work this is the wrong instrument, full stop. Where the value of the footage is that it records something, generation cannot supply that property, whatever it supplies instead.

Common questions

Why not generate a long sequence directly?

Because output arrives as short shots, and a single continuous take of anything complicated is the hardest version of the job even with a camera. Sequences are assembled in an editor from separately generated shots, which is how film has always worked.

How is continuity between shots maintained?

By fixing what you can outside the generation — identical wording for subject and location, a stated light direction, reused seed images where available — and by planning cuts that tolerate drift, such as angle changes and cutaways. Needing fewer matched shots is the strongest fix.

Why does planning matter more for motion than for stills?

Because a clip cannot be locally repaired. Any change means regenerating the whole shot, so decisions about framing, aspect ratio and treatment have to be made before generating rather than adjusted afterwards.

Images & Audiovideocontinuityeditingplanning
Devika Menon
Reporter, Prompt After Prompt

Devika has been reporting on prompt craft, writing with ai, images & audio since long before it was fashionable and reads the small print so you do not have to.