Why identity drifts
The model has no stored representation of your character. Every generation is a fresh sample conditioned on whatever you supplied. Two generations from the same description land on two nearby but different points, and nearby is not the same.
Faces are where this becomes intolerable because human vision is specialised for faces. A three per cent change in the distance between the eyes reads as a different person. A thirty per cent change in the shape of a chair reads as the same chair. The model is no worse at faces than at furniture; you are far better at noticing.
The four mechanisms, in order of strength
- A fixed first frame. The strongest thing available, and free. If shot two begins from a still containing the face you want, the face survives that shot. This is the single biggest argument for the still-first workflow from the earlier lesson.
- Reference or element features. Kling's elements, Runway's references, Higgsfield's character tools, Sora's consent-gated cameos. You supply images of a subject and the model conditions on them. These give you a consistent type of person and a consistent wardrobe. They do not reliably give you the same actor across ten shots.
- A trained character model. A LoRA on an open-weights base is the strongest identity lock there is. It costs you a dataset of fifteen to thirty varied images, a training run, and the local setup from the last lesson. See fine-tuning.
- Description alone. The weakest. "A woman in her forties with short grey hair" produces a genre of woman, differently each time.
The workflow that holds up
Build a character sheet before you generate a single clip:
- The face at frontal, three-quarter and profile.
- Full body in the exact wardrobe, front and back.
- The same lighting direction in every one of them.
Then every shot's first frame is built from that sheet — generated with the reference where the tool supports it, and composited by hand in Photopea or GIMP where it does not. Head from one image, body from another, background from a third. Compositing feels like defeat the first time you do it. It is faster than the twelfth generation, and it is deterministic, which no amount of re-rolling is.
Lighting and wardrobe break the cut before the face does
This is the part almost nobody warns you about. Two shots can have a genuinely matching face and still refuse to cut together, because the key light is on the left in one and the right in the other, or the shirt is a slightly different blue, or the sun is at a different height.
Editors call this continuity and audiences read it instantly without being able to name it. The feeling is "something is off", and people usually blame the face when the face was fine.
So decide the light direction for the scene before you start, put it in every first frame, and check the wardrobe colour by putting the stills side by side rather than by remembering.
Colour drift between generations can be pulled together afterwards. In DaVinci Resolve's free tier, put two clips on the timeline, use the shot-match to move one toward the other, then correct the skin tone by eye. That fixes a warm-versus-cool mismatch. It does not fix a different jacket.
Shoot around the problem
The professional answer is not to solve identity. It is to need less of it. Design the sequence so the face carries less of the load:
- Over-the-shoulder and back-of-head shots.
- Hands, feet, objects, details, the thing being looked at.
- Wide shots where the face is forty pixels across.
- Reaction cut to something other than the character.
- Fewer shots of the character than your instinct says.
A three-shot scene where the face is clearly visible once is far more achievable than a six-shot scene where it is visible every time. It is also, independently, better editing — which is the pattern this whole subject keeps producing.
The honest number
Even with references, a character sheet and matched lighting, expect to reject a substantial share of generations for identity alone, and expect the odds to compound across a sequence: a per-shot success rate that feels acceptable becomes a low probability of getting eight consecutive shots that all match.
Anyone showing you a character-consistent, ten-shot narrative without mentioning an enormous discard pile is showing you the survivors and not the attempt.
Today: build a character sheet for one person — three angles, one wardrobe, one light direction — before you generate any motion at all.
Before you move on