Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Most of your audience has the sound off

Editing Video · lesson 7 of 9 · 9 min

The default is silence

Video in a feed autoplays muted. People watch on a bus, in a shared room, in a lecture, next to a sleeping child. A large fraction of the views on anything you post will begin, and often end, without sound.

That is before you count the people who need captions rather than prefer them: deaf and hard-of-hearing viewers, and — very relevant to this classroom — the large number of people watching in their second or third language, who read English considerably faster than they parse spoken English at conversational speed with an unfamiliar accent.

Captions are not an accessibility box to tick at the end. On short-form video they are the primary channel.

Two kinds of caption, and you often want both

Burnt-in (open) captions are rendered into the pixels. They always appear, on any platform, in any player, in autoplay-muted feeds. They cannot be turned off, cannot be translated, cannot be read by a screen reader or indexed by a search engine, and small thin text can be mangled by aggressive compression.

Sidecar subtitles are a separate file — .srt or .vtt — that the player overlays. They can be switched off, restyled, auto-translated, indexed by search, read by assistive technology, and used to generate a transcript. They do not exist in a muted autoplay preview unless the platform chooses to show them.

For a vertical short: burn them in and upload the sidecar file. For a long-form YouTube video or a course: sidecar, always, and burn in only key phrases. Every editor in this course exports SRT, including Kdenlive and Shotcut.

Automatic captions and how they fail

Speech recognition is genuinely good now. On clear speech, in a well-represented accent, with common vocabulary, expect somewhere in the region of 90 to 95% word accuracy.

Accuracy falls off sharply on exactly the things this audience deals with daily: non-native and regionally accented English, code-switching mid-sentence between English and another language, proper nouns, place names, technical terms and acronyms.

And the failure mode is the dangerous part. Recognition does not produce gibberish that you notice. It produces confident, grammatical, wrong sentences that read perfectly and say something else. "Deploy the model" becomes "destroy the model". Nobody proofreads a caption track that scans fine.

So: never publish an auto-caption track unread. Reading a five-minute track takes about four minutes and it is not optional.

Free ways to generate them: YouTube's own automatic captions, which you then edit in Studio; Whisper, an open-weights speech model that runs locally — whisper.cpp will transcribe on an ordinary laptop CPU with no GPU and no upload, which matters if the audio is confidential; and various free web transcribers, with the obvious caveat that you are uploading the audio to someone else. Resolve has transcription built in on some versions and tiers, so check yours.

Making captions legible

The specifics matter more than the styling.

  • Two lines maximum, roughly 37 to 42 characters per line. Longer lines force the eye to travel and people lose the picture.
  • Reading speed around 15 to 17 characters per second is comfortable. Above about 20 people stop reading and start skimming.
  • Minimum time on screen of about a second, even for a single word, or it flashes.
  • Break lines at grammatical joins. Break before a preposition or a conjunction, not in the middle of a noun phrase. "The reason the export / failed" reads badly; "The reason / the export failed" reads well.
  • Contrast has to survive a changing background. Text over video sits on a different colour every frame. Use a semi-opaque box behind it, or a solid outline, or a tight drop shadow. Plain white text with nothing behind it will disappear over a bright sky.
  • Never use a hairline or thin weight. At phone size, after platform compression, thin strokes break up. Medium or bold, always.
  • Stay out of the interface. On vertical short-form, the bottom fifth or so of the frame is covered by the username, description and buttons, and the top by the status bar and platform chrome. Put captions in the middle third. Editors lose text to this constantly because the editing canvas has no interface on it.

Titles

The same discipline as Design Foundations, with two additions specific to video.

Keep text inside a safe margin — roughly 5% in from every edge — because players, televisions and platform crops all eat the outer band. And give titles enough time: the honest measure is that you should be able to read the text aloud twice at a normal pace while it is on screen. You will always underestimate this, because you wrote it and you already know what it says.

Resist animating for the sake of it. A title that slides, spins, bounces and glows has spent the viewer's attention on the animation rather than the words.

Free tools

Resolve has a full subtitle track and exports SRT. Kdenlive has a decent subtitle editor. CapCut generates captions automatically and styles them well, which is a large part of why it dominates short-form editing. Subtitle Edit is free, open source, and better at the fiddly work of timing and line-breaking than any video editor. On a phone, CapCut and most mobile editors will auto-caption and burn in; proofread on the phone before exporting.

Today

Take a video you have already published, read its automatic caption track end to end, and count the wrong words. That number is what your silent viewers received.

Before you move on

A creator burns captions into a vertical video and positions them near the bottom of the frame, where they looked balanced on the editing canvas. On the platform, viewers say the captions are unreadable. What went wrong?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly