Hands doing precise tasks
Threading a needle, tying a lace, typing, playing a guitar, counting notes. The mechanism is contact and occlusion. The model learned the statistics of what hands look like; it did not learn the constraint that a finger cannot pass through a string. Contact events are exactly where a plausible-looking frame becomes an impossible one, and hands make contact constantly, at speed, while occluding each other and the object.
This is not a resolution problem and more sampling steps will not fix it.
Work around it: cut away at the contact. Hand approaches, cut, result. A century of film grammar exists for precisely this reason, because real crews could not always film the contact either.
Readable text
Signs, packaging, screens, anything with letters in the frame. Text is a symbolic system with a small alphabet and strict rules; the model has pixel statistics. Image models improved on this enormously. Video models are behind, and there is an extra reason: a letterform must stay identical across a hundred and fifty frames, and one that shifts by two pixels a frame reads as melting even when each individual frame is fine.
Work around it: composite the text in afterwards, in the editor, as an actual title with an actual font. Or put it in the first frame and keep the camera nearly still — which sometimes survives and usually does not.
Physical cause and effect
A glass that falls and shatters into the right number of pieces. A ball that bounces lower each time. Cloth that settles and stays settled. Fire that reacts when the wind changes. The model interpolates plausible-looking frames; it does not simulate. There is no state, no mass, no conservation of anything.
It gets short, common, heavily represented physics roughly right — a person walking, water running downhill, hair moving in a breeze — because those patterns are everywhere in the training data. Anything unusual it gets wrong, confidently.
Work around it: single events, short, and never in slow motion. Slow motion is where fake physics is most visible, because you have given the viewer time to check.
A specific real product
Your client's bottle, with their label, their cap, their exact proportions. The model has no reference for that object. It produces a member of the category with a label that is almost right — and almost right is worse than obviously wrong, because it reads as a counterfeit rather than as a stylisation.
Work around it: this is a compositing job, not a generation job. Generate the environment and the camera move; mask and track the real product in. Or shoot the product on a phone against a plain wall, which is one of the places where thirty seconds of real filming beats an afternoon of generating and everyone can see it did.
Continuity over minutes
Covered earlier, but the mechanism is worth stating once more because it explains four of the other five items. There is no persistent object memory and no scene model. Every frame is inferred, the only anchor is the conditioning image, and error grows with distance from that anchor. Anything that requires the same object to be the same object several shots later is fighting the architecture, not the settings.
Action timed to a beat
"She looks up on the fourth beat" is not currently something you can ask for. Duration maps loosely onto the request and the timing of an action within a clip is not controllable.
Work around it: generate longer than you need and find the frame in the edit, then slip the clip until the action lands. That is what an editor does with real footage anyway.
Two things people over-claim in both directions
Not all of this gets fixed by the next model. Hands and short-range physics have improved measurably from one generation of models to the next, and will keep improving. A specific real product is different in kind — it is not a capability gap, the information simply is not available to the model, and no scale of training changes that.
Not all of it is permanent either. Text in image models was a standing joke in 2022 and mostly works now. So write down the mechanism rather than the verdict, and re-test this list yourself every few months on your own shot rather than believing anyone, including this lesson.
The most valuable thing you can tell a client or a team is where the tool stops. Being the person who says on day one that a shot is not achievable is worth considerably more than being the person who can drive the tool and discovers it on day fourteen with the deadline in sight.
Before you move on