Directing motion, and what actually lands
Why some camera words work
The mechanism is the same as for image prompts. Video training data is captioned, increasingly by machine, and those captions describe motion using whatever vocabulary the captioner used. Terms that appear consistently in that corpus behave like reliable controls. Terms that do not are decoration.
In practice this means standard film vocabulary works, because it is what any competent captioner uses and what stock video libraries have always tagged with:
slow push in,dolly out,pan left,tilt up,crane downhandheld,locked off,tracking shot,orbitrack focus,shallow depth of fieldslow motion,time-lapse
And invented or compound descriptions do not: the camera swoops dramatically around the subject and then rises to reveal the city is one caption nobody wrote, and the model will pick whichever fragment it recognises.
The practical instruction is to learn a small amount of real film vocabulary and use it plainly, one move per shot. This is a genuine reason to read a cinematography primer: the words are the interface.
Subject motion versus camera motion
These are separate and models frequently confuse them. moving left is ambiguous — the subject, or the frame? Written captions resolve this by convention and the conventions are inconsistent.
Say which. The camera pans left while the woman stands still is unambiguous and lands far more often than panning left.
A useful diagnostic: if you cannot tell from your own prompt whether the background should move, neither can the model.
Explicit trajectory conditioning
Several systems now accept camera motion as data rather than words — a path, a set of positions per frame, or a drag gesture indicating where an element should travel. Where this exists it is categorically better than prompting, for the same reason a depth map beats "on the left": it is spatially explicit and does not depend on caption vocabulary.
Where it does not exist, the free substitute is video-to-video from a source clip with the camera move you want. That source can be a phone video, or a Blender render of a moving camera over grey blocks, which takes ten minutes to set up and gives exact, repeatable moves.
The motion-strength setting
Most image-to-video systems expose an amount-of-motion control. It is worth understanding what it trades.
Low motion gives you stable, coherent, nearly-still clips. It is also where models look best, and where the demonstration reels quietly live: a slow push on a static subject hides almost every temporal weakness.
High motion gives you movement and buys the whole class of failures — morphing, limbs multiplying, backgrounds reorganising. Fast motion is where the temporal compression from the first lesson smears, and where the attention window is stretched hardest.
The professional consequence: design shots that need little motion. This sounds like a constraint and it is also good practice — a held shot with a small camera move and one thing happening is a better shot than a busy one in any medium. Most people arrive wanting the model to do something spectacular and get much better results the moment they ask it for something restrained.
What no model does yet
Two things worth knowing before promising them.
Continuity of action across a cut. Two shots of the same action from different angles, matching. This is basic film grammar and there is no reliable route to it in generation. The workaround is to generate one wide take and crop into it for the closer shots, which is how a lot of low-budget live action is cut anyway.
Precise timing. Making something happen on a specific frame — a hit on a beat, a door closing on a word — is not controllable. You get the action somewhere in the clip and you place the cut in the edit. Budget for it.
There is also a cheap substitute for a camera move that nobody should be embarrassed to use. Generate a wider, static shot at higher resolution, then create the move in the edit by animating a crop across it. A push in, a slow pan, a reframe — all of these are a keyframed crop, they are exact, they are repeatable, and they introduce no temporal artefacts at all because nothing in the picture is changing. Documentary editors have made a living from this technique on still photographs for decades. Against a generated camera move that lands one time in five, it is frequently the better shot as well as the cheaper one.
The honest summary: camera control has improved from "unusable" to "usable within film vocabulary", and precise directing is still done by conditioning on real footage or a 3D render rather than by asking.
The one thing to keep
Camera and motion instructions work only where the training captions contained consistent film vocabulary, so standard shot language lands and invented descriptions do not, and explicit trajectory conditioning beats both.
Before you move on
Which prompt is most likely to produce the intended shot on a current video model?
Pick the one you would defend. Nobody sees your answer.