Conditioning a clip: stills, frames and video in
Text alone is the weakest way to make a clip
Prompting a video model with text only asks it to decide the subject, the composition, the lighting, the camera and the motion, all at once, from a sentence. Every uncertainty in the image case is present, plus motion.
The result is the familiar experience: attractive clips that are not what you asked for, at a hit rate that makes any planned sequence impractical.
Every other conditioning route is better, and knowing what each one pins is the whole skill.
Image to video
Supply a still as the first frame. The model continues it.
This is the dominant professional workflow and the reason is simple: it moves the entire composition problem into the image domain, where you have masks, structural conditioning, fine-tunes, compositing and unlimited cheap iterations. You settle the frame using everything in the previous module, then ask only for motion.
What it pins: composition, subject appearance, colour, lighting, style at frame one.
What it does not pin: what happens next. The subject can still turn into someone else by second four. Motion is still decided by the model.
A practical detail people miss: the input still is often re-encoded, so the first frame of the output is not pixel-identical to your input. For a shot that must cut cleanly from a still, check this and, if needed, replace frame one in the edit.
First and last frame
Supply both ends. The model interpolates a plausible path between them.
This is the strongest control available for a specific action, and it is under-used. It turns "make him pick up the cup" into "here he is with an empty hand, here he is holding the cup", which is a much better-posed question. It is also how you make a clip loop: use the same image for both ends.
Its failure is characteristic and worth expecting: given two ends that cannot be connected by a plausible continuous motion, the model produces a morph rather than a movement. Objects melt into each other. When you see morphing, the two frames were too far apart, and the answer is an intermediate frame rather than a better prompt.
Video to video
Supply a clip; get a restyled clip. Underneath this is usually per-frame structural conditioning — depth or edges extracted from each input frame — plus the model's temporal layers.
What it pins: motion and geometry, exactly, because they come from the source.
What it does not pin: appearance stability, which is why early attempts at this flickered so badly. Current models are much better, and the remaining tell is texture that boils on surfaces the source video kept flat.
This is the route for anyone who can shoot. A phone video of the actual motion, restyled, beats generated motion for anything requiring specific action. It also solves the rights position for the performance, which the talking-heads lesson takes up.
Reference conditioning
Supply an image not as a frame but as a description of who or what should appear. Identity and style come from the reference; the shot is generated freshly.
Useful for keeping a character across shots that cannot be one continuous take. Weaker than a first frame, more flexible.
Choosing, in one line each
- Need a specific composition: first frame.
- Need a specific action: first and last frame, or video to video.
- Need the same character across several shots: reference conditioning, plus a face pass afterwards.
- Need a specific camera move: video to video from a phone clip or a Blender render.
- Need something to look nice and do not much mind what happens: text only.
There is one more route worth knowing, because it is free and often overlooked: animating a still without a video model at all. A slow push, a parallax move on a layered image, a subtle drift of clouds, a flicker of light — all of these are keyframe work in an ordinary editor, and they cost nothing, never flicker, never drift, and are perfectly reproducible. A great deal of what people generate as video would be better made as a still with a move on it. Kdenlive, Shotcut and Resolve's free edition all do this; so does Blender. Ask whether the shot actually needs generated motion before spending generated motion on it.
The honest limitation across all of them: conditioning constrains the frames you supply and the frames near them. Nothing constrains the middle of a long clip. This is the same window limit as the previous lesson, and it is why the answer to almost every video problem in practice is "make the shot shorter and control both ends".
The one thing to keep
A video model can be anchored by a first frame, a last frame, a full input video or a reference image, and each conditioning route constrains a different thing and fails differently.
Before you move on
A first-and-last-frame generation returns a smooth morph rather than a movement, with objects melting into one another. What is the likely cause?
Pick the one you would defend. Nobody sees your answer.