Talking heads, lip sync and the consent line
How the mechanism works
There are two broad approaches and both are now cheap enough to run on ordinary hardware.
Audio-driven reenactment. A model takes a still photograph or short clip of a face and an audio track. It predicts, per frame, the mouth shape and head motion that would produce that audio, and renders the face accordingly. Underneath, the audio is converted to a sequence of features that correlate with mouth position, and the face model is conditioned on them frame by frame. The rest of the image is held.
Full generation with audio conditioning. A video model generates the whole shot with the audio as a conditioning input, producing head, body and background together.
The first is more stable and more obviously an edit. The second is more convincing and drifts more.
The quality determinant most people miss is the audio, not the face. Clean, close-miked speech produces accurate mouth shapes; noisy or compressed audio produces mush, because the features the model reads are degraded before it starts.
What it is genuinely good for
This is not a purely dangerous technology and it is worth being specific about the useful cases, because a blanket warning teaches nothing.
- Dubbing with matched lips. A film translated into another language, with the mouths adjusted to the new dialogue. This is a real improvement in accessibility and the licensing arrangements for it are ordinary.
- Presenter video at scale for training material, where a real presenter has consented to a synthetic version of themselves.
- Restoring a damaged performance in post-production — a line re-recorded and re-synced rather than reshot.
- Accessibility, including giving a synthetic presence to someone who cannot appear on camera.
Every one of these has the same structure: the person depicted agreed, specifically, in advance.
The consent standard
The voice module in this course sets out the standard in full and it applies identically to faces. In short, consent must be:
- Informed — the person knows what will be produced and in what contexts.
- Specific — to this use, not to any future use anyone thinks of.
- Written — because the dispute will be about what was agreed.
- Revocable — with a stated process and a stated timescale.
- Compensated, where the use is commercial.
And separately from consent, the audience needs to know. Someone who consented to a synthetic presenter has not thereby consented on behalf of the viewer to be deceived. Disclosure is a second duty, and in several jurisdictions it is now a legal one, which the final module sets out.
The part that is not solvable by good practice
Every technique described here works from a photograph and an audio sample. Both are publicly available for most people who have ever appeared online. That means the capability to produce convincing video of an arbitrary person is widely distributed and cannot be recalled.
The harms are documented and heavily skewed. The overwhelming majority of synthetic video circulating is non-consensual sexual imagery, and the overwhelming majority of its targets are women. Fraud is the second category — a video call from a colleague's face authorising a payment. Political fabrication is real and, by volume, third.
The realistic responses are institutional rather than technical:
- Verification out of band. A second channel for any instruction that moves money or grants access. This defeats the fraud case entirely and costs nothing to adopt.
- Provenance, covered in a later module — signed capture rather than detection after the fact.
- Law. Several jurisdictions have created specific offences for non-consensual synthetic sexual imagery and for fraudulent impersonation; the coverage is uneven and expanding.
What does not work is detection by eye. Assume that any video of a person can be fabricated and that the question worth asking is where the file came from, not whether the mouth looks right.
One practical note for anyone commissioning this work. A performer's agreement to appear on camera is not an agreement to a synthetic version of themselves, and a standard release form written before any of this existed does not cover it. Several performers' unions negotiated specific terms on exactly this point, and the shape they arrived at is a reasonable template even outside a union context: separate consent for the creation of a digital replica, separate consent for each use, a stated term after which it expires, and payment for use rather than only for the capture session. If you are the performer, ask for those four things. If you are commissioning, offering them unprompted is cheap and it is the difference between a working relationship and a dispute.
For your own work the line is simple enough to state in a sentence: do not make a recognisable person say or do anything without their specific written agreement, and tell the audience when what they are watching is synthetic. Both parts are required. Neither is difficult.
The one thing to keep
A talking-head system drives face and mouth motion from an audio track, which makes convincing video of a person saying words they never said cheap, and the only real control on it is consent recorded before the fact.
Before you move on
An organisation wants to use a synthetic version of a consenting presenter for internal training videos. What else is required beyond their consent?
Pick the one you would defend. Nobody sees your answer.