Two families of talking-head tool
Audio-driven lip-sync. You supply a video or a still plus an audio track, and the model reshapes the mouth region to match the phonemes. The older open-source line — Wav2Lip and its descendants, LatentSync and similar — does the mouth only, and does it well enough at small sizes. Hosted products such as Hedra, HeyGen and Synthesia, and the sync features inside the larger platforms, animate the mouth plus some head movement, blinks and eye direction.
Performance transfer. You record yourself acting on a webcam and the model maps that performance onto a generated character. Runway's Act-One is the best-known example of the class. This is substantially better than lip-sync alone, because the timing, the pauses and the eyebrows come from a person who understood the line. Delivery is most of acting, and no phoneme model invents delivery.
Where they break, and why
- Profile angles. The mouth has to be visible and roughly frontal. Past about thirty degrees off-axis the model is guessing at geometry it cannot see, and the jaw starts sliding. Keep the head near frontal for any shot that has to speak.
- Plosives. B, P and M require the lips to close completely. Many models under-close them, and the result reads as badly dubbed foreign footage — most viewers cannot name what is wrong but all of them notice. Slowing the delivery slightly improves it.
- Teeth and tongue. The interior of the mouth is still where close-ups fall apart. Stay at a medium shot and the problem largely disappears.
- Long takes. The head loses its anchor over time, exactly as everything else does. Cut every few seconds.
Voice
Cloning a voice from a short sample is a solved problem, and it is cheap. The free path is real: open models such as Coqui XTTS, Piper and Kokoro run on modest hardware and produce serviceable synthetic speech. Hosted services like ElevenLabs are better and cost money. Audacity — free, runs on almost anything, including very old machines — is enough to trim, level, remove a hum and normalise. Recording a real voice on a phone in a quiet room, cleaned in Audacity, still beats most synthetic speech for anything longer than a few sentences.
The line, stated plainly
You need permission to use a real person's face or voice. This is not a preference, a style question or a thing that becomes acceptable at sufficient production value.
- Get it in writing, and get it specific: what the likeness will be used for, on what platforms, for how long, and whether it may be altered or made to say new things.
- Consent to appear is not consent to be synthesised. A client who filmed an interview with an employee last year did not thereby acquire the right to generate new sentences in that employee's voice. Those are different acts and the first does not license the second.
- Public figures and dead people are not a loophole. India's courts have issued a run of personality-rights orders since 2023 protecting named actors' likeness and voice. Denmark moved to give people a right in their own likeness. Many US states have publicity-rights statutes, and the 2025 TAKE IT DOWN Act obliges platforms to remove non-consensual intimate imagery, synthetic included. The direction of travel is one way.
- The consent mechanism inside some products — Sora requires a person to record a verification video before their likeness can be used as a cameo — is a decent model for how to think about this. It is not a substitute for your own paperwork on a commercial job.
That is the shape of it, not advice for your case. A lawyer in your own country is the person who answers it for a specific job, and the answer differs by country more than people expect.
The labelling duty, which is separate
Consent is about permission. Labelling is about the audience, and it is a distinct obligation that applies even when everyone involved agreed.
Several countries now require synthetic media to be labelled visibly. China's rules require both a label a person can see or hear and an implicit label inside the file's metadata. India amended its IT Rules for synthetically generated information in a way that prescribes visibility — a label covering a defined portion of the frame and the opening portion of the audio, plus a duty on platforms to verify declarations. The EU AI Act carries a transparency obligation covering deepfake disclosure. Timetables have shifted more than once, so confirm the current position rather than trusting any date, including this one.
Content Credentials, built on the C2PA standard, and invisible watermarks such as SynthID are real and worth keeping in the file. Both are fragile. A screenshot destroys them. Many social re-encodes strip them. The label that survives contact with the internet is the one burned into the frame.
So put it in the video, at the start, where a viewer sees it before they believe it. It costs you a title card and about two seconds.
Today: if you have a project with a real person in it, write down four things — whose face, whose voice, what you hold in writing, and what the video says on screen about being generated. Any blank is your next task, before the next generation.
Before you move on