Where the machines actually help
Strip out the marketing and four things have genuinely changed in this workflow.
Speech to text is the dull one and the most useful. Whisper, released with open weights, transcribes and time-stamps audio, and runs offline on an ordinary laptop through whisper.cpp on modest hardware. That gives you captions, a searchable transcript, and subtitle files, free. Captions raise completion rates on muted feeds and, in several contexts, are an accessibility obligation rather than a nicety. Auto captions get proper nouns, place names and technical terms wrong every single time. Read them before you burn them in.
Transcript-based editing — delete a sentence in the text and the audio cuts with it — is a real speed change for interviews and talking-head work.
Generated music. Suno and its competitors will produce a fifteen-second bed that most listeners cannot distinguish from a library track.
Text to speech and voice cloning. ElevenLabs is the reference commercial product. Open alternatives include Piper, which is small and fast enough to run on a Raspberry Pi, and Kokoro, which is small and permissively licensed, alongside a rotating cast of larger models. Check each model's licence individually: several well-known open voice models are released under terms that forbid commercial use, and "the weights are downloadable" is not the same as "you may sell work made with it".
Enhancement models for de-noise, de-reverb and studio-voice processing are much better than the old signal-processing versions, and fail in the same direction: pushed hard, they invent detail. On a heavily processed voice you can hear consonants being reconstructed slightly wrong.
Consent, stated plainly
A voice is a person. Cloning one without agreement is not a technique question, and there is no version of it that is a style choice.
Get the agreement in writing, and make it specific:
- For which project.
- For which uses — this video, or anything the company publishes.
- For how long.
- Whether the trained model itself is kept afterwards, and who can run it.
"You can use my voice for this video" is not permission to hold a voice model on a server for two years. The gap between those two things is where most of the real disputes sit.
Read what the service does with the samples you upload, including for your own voice. And note the second-order risk: a cloned voice on a phone call has already been used to extract money from families and payments from finance departments. Publishing long, clean, isolated samples of a colleague's voice makes that easier. Not a reason to stop working — a reason to think about what you post.
The law is moving in one direction here. Denmark moved to give people a right in their own likeness and voice. Several US states have right-of-publicity statutes that already reach a deliberate soundalike. The 2025 TAKE IT DOWN Act obliges US platforms to remove non-consensual intimate imagery including synthetic material.
Disclosure and ownership
Disclosure
Several countries now require synthetic media to be labelled visibly, not merely tagged in metadata.
India's amended IT Rules for synthetically generated information prescribe visibility, and for audio specifically a label in the opening portion of the audio. China's rules, in force since September 2025, require both a label a person can see or hear and one embedded in the file's metadata. The EU AI Act's transparency provisions require machine-readable marking of generated output and disclosure of deepfakes, on a timetable that has been amended more than once — check the current position rather than trusting any date you read, including this one.
The rule that survives all of these statutes: if a listener could reasonably take the voice for a real person speaking, say that it is not, at the start, in the audio, where they hear it before they believe it. Not in the description. Not in the credits.
Ownership
Generated music sits in the same unsettled place as generated images. In the United States a prompt alone is unlikely to make you an author; your own arrangement, your recorded parts and your edits are yours. Other countries have reached different conclusions on similar facts. Making things with AI sets out the country-by-country shape of it.
Separately from ownership, the training-data question is live: the major labels sued the generated-music companies in 2024, and licensing arrangements have been reported since. That is the shape of it. A lawyer in your country is the person who answers it for your case.
What has actually changed in the work
A one-person team can now caption a video, transcribe an hour of interview, produce a music bed and lay a temporary voiceover in an afternoon — work that used to mean booking people and a room.
What has not changed is the part that was ever the job: deciding what to cut, what the piece is about, whether the music fits the point, and whether the reader can follow it. A generated voice is good at neutral information delivery. It is still audibly wrong at sarcasm, grief and comic timing, because those live in the timing between words, and a real performance still wins there by a distance.
One habit
If you use a generated voice in anything you publish this week, write the disclosure line before you generate the audio, not after you finish the edit. Written first, it is one sentence. Written last, it is a change to a locked timeline, and it does not get made.
Before you move on