The asymmetry nobody warns you about
A soft, badly lit, wobbly picture with clean sound reads as low-budget but competent. People watch it to the end. A sharp, well-composed picture with hissy, echoing, uneven sound reads as broken, and people leave in the first fifteen seconds.
This is not taste. Speech intelligibility is effortful: if a listener has to reconstruct words from a noisy signal, they spend attention on decoding rather than on what you are saying, and after twenty seconds of that they stop. Nothing in the picture compensates, because nothing in the picture is doing that work. If time on a project is short, spend it on the audio.
The numbers that actually matter
Digital audio is measured in dBFS — decibels relative to full scale — counting down from zero. Zero is a hard ceiling, not a warning: above it, samples are flattened, the waveform develops a square edge, and it crackles. Unlike analogue tape there is no warmth up there and nothing to recover. Clipping is permanent damage.
Working targets for spoken word:
- Dialogue peaks around -6 dBFS, average level around -12 dBFS. That leaves headroom for a laugh or a raised voice.
- Music sitting under dialogue: roughly -18 to -24 dBFS, quieter than feels right to you, because you already know the words.
- The final mix is normalised for the platform at export.
The two commonest phone-video faults are opposite versions of one mistake: gain at maximum so everything clips, or recording so quietly that you add 20 dB in post — which raises the hiss and the traffic and the fridge by 20 dB too. The noise floor is baked in at the moment of recording.
Distance beats equipment
If you buy one thing, buy a cheap wired lavalier microphone — around ₹800 to ₹1,500, or ten to twenty dollars — and clip it about 20 to 30 centimetres from the mouth, out of frame. It will beat an expensive microphone two metres away, every time.
The reason is the ratio of direct sound to reflected sound. Sound reaching the mic straight from the mouth falls off quickly with distance; sound bouncing off walls and ceiling is spread evenly through the room and barely falls off at all. Halve the distance and you dramatically improve that ratio. This is why phone video shot across a room sounds like a bathroom: not because the phone mic is bad, but because most of what reaches it has been off a wall first.
Then turn off the fan, the air conditioner and the fridge, close the window, and put a blanket behind the camera if the room is bare. None of it costs anything.
Room tone, and the holes in your dialogue
Record thirty seconds of the room doing nothing at every location. Everyone forgets, and everyone regrets it.
Every room has a noise floor — air, hum, distant traffic. When you cut dialogue, that floor stops and restarts at every join, and the ear hears the sudden silence as a hole punched in the scene. It is one of the clearest amateur tells there is. Lay room tone on a track under the whole scene, fill any true gaps with it, and the holes vanish.
While you are there: a high-pass filter at 80–100 Hz on every voice track removes rumble, handling noise and air conditioning without touching the speech.
Noise reduction, and exactly how it fails
Classic noise reduction works by spectral subtraction. You give it a sample of noise alone, it builds a profile of how much energy sits in each frequency band, and it subtracts that amount everywhere.
Push it too far and quiet harmonics of the voice fall below the estimate, so they are removed in some frames and kept in others. Those fragments flutter in and out and you get the watery, swirling, underwater artefact engineers call musical noise.
The rule: if you can hear the noise reduction, there is too much of it. Six decibels that leaves a little honest hiss beats twenty that leaves a robot. Audiences forgive background noise. They do not forgive artefacts, because artefacts sound like something is wrong with the file.
Free tools: Audacity's noise reduction, and the EQ and dynamics on Resolve's Fairlight page — though some Resolve restoration plug-ins are Studio-only.
The newer speech tools, and where the line is
A different class of tool does not subtract noise. It runs the recording through a model and resynthesises the voice. The results can be startling — a phone recording from across a room comes back close to studio.
Two cautions. It is generating audio, so it can shift timbre and invent detail that was never there, which makes it inappropriate wherever the recording is evidence, journalism, or a record of what someone said. And once a tool goes past cleanup into replacing words, changing accents or producing a voice, that is synthetic media: it needs a visible label — several countries, India among them, now require labelling of synthetic content — and somebody's voice needs their written consent. Consent for a likeness or a voice is not a style question. Making Things With AI covers where that stands.
Two traps that catch everyone once
- Mono recorded into one channel. A lavalier plugged into a single input often lands on the left channel only, so anyone listening on one earbud hears silence. Set the clip's audio to mono, or copy it to both channels.
- No sync reference. If picture and sound came from different devices, clap once, hard, in frame at the top of every take. A clap gives you a single unmistakable spike to line up. Waveform auto-sync works well but fails on noisy or quiet takes, and then you will want that clap.
Today
Play something you made with the picture turned off. If you would not listen to it as a podcast, the picture is not saving it.
Before you move on