Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Voice, and the line you do not cross

Making Things With AI · lesson 6 of 8 · 9 min

Cloning is fast now

Modern voice models do not retrain on your voice. They extract a speaker embedding — a compact numerical description of timbre and delivery — from a short reference clip, then condition a speech model on it. Three to thirty seconds is usually enough.

Quality of reference beats quantity. One clean, dry, mono recording with no music, no reverb and no second speaker will beat ten minutes of a noisy phone call. Match the delivery too: a clone built from calm reading sounds wrong shouting.

What this is genuinely good for

Audiobook narration in an author's own voice. Dubbing a course into eight languages while keeping the teacher's voice. Voice banking for people losing speech to ALS or throat cancer, which on its own justifies the field. Fixing one mispronounced word in an hour of finished narration without recalling the talent.

Every one of those shares a feature: the person whose voice it is agreed.

The line

You may clone a voice when the person it belongs to has agreed, knows what it will be used for, and can withdraw. That is the whole rule.

It is not softened by any of the following, all of which people say with real confidence:

  • It is satire.
  • He is a public figure.
  • It is my father and he would have found it funny.
  • I paid for a dataset of their podcast.
  • I put a disclaimer in the description.

What ignoring it enables

Family-emergency fraud. A cloned voice of someone's child on a deliberately bad line, asking for money urgently. It works because fear disables judgement before scepticism arrives. The defence that actually works is a family passphrase agreed in advance, offline, and never sent over anything.

Payment fraud, using an executive's voice lifted from a conference recording, aimed at a finance team under time pressure.

Elections. In January 2024, voters in New Hampshire received a robocall in a synthetic voice imitating the US president, telling them not to vote in the primary. The US telecoms regulator then confirmed that AI-generated voices in robocalls fall under existing anti-robocall law, and proposed a multi-million-dollar fine against the operative behind it.

The law is arriving, unevenly

Tennessee's ELVIS Act, in force since 2024, added voice explicitly to the state's right of publicity — a music state protecting its industry.

India has built protection through personality-rights injunctions rather than a statute. The Delhi High Court protected Anil Kapoor's name, image, voice and mannerisms in 2023, and in 2024 the Bombay High Court granted the singer Arijit Singh an order dealing directly with AI voice cloning.

Denmark moved to give individuals a copyright-like right in their own likeness and voice.

The EU AI Act requires anyone deploying a deepfake to disclose it.

A US federal statute on voice and likeness has been introduced more than once and, as far as I know, has not been enacted. Check the current position rather than trusting that sentence.

The direction is unambiguous even where the text is not final. This is becoming a right people hold in themselves.

Consent that actually holds

If you commission a voice, get a written agreement naming:

  • whose voice it is, and which recordings the model was built from
  • what it may say, as categories rather than vibes
  • where it may appear, and for how long
  • what it may never do: political speech, endorsements, sexual content, anything defamatory
  • how the person withdraws, and what happens to the model files when they do
  • what they are paid, including for reuse

Keep the recording of the consent conversation with the model files. And say the sentence out loud once, in the room: we are making a copy of your voice that can say things you never said. If that makes the conversation awkward, the awkwardness is information.

Dead people

Estates can license, and often do. The harder question is not legal. In a 2021 documentary about the chef Anthony Bourdain, the director had a few lines of Bourdain's own written words performed in a synthetic version of his voice, and did not tell the audience. The words were his. The performance was not. Viewers felt deceived, and the fallout has shaped practice since.

The lesson holds generally: disclosure has to reach the audience before the belief does. Not in the credits.

Before you move on

A documentary team clones the voice of a subject who has died, to read letters the subject genuinely wrote. The family has agreed and been paid. A note appears in the closing credits. What is the sharpest criticism?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly