A voice is the biggest single upgrade available to a faceless video and the one most people skip, usually because they assume it requires equipment or an on-camera presence. It requires neither.
What it actually buys
- Watch time. A viewer who is listening stays longer than one who is reading, and watch time is the metric doing most of the work — reel analytics: the four numbers worth reading.
- Searchability. Spoken words are transcribed by the platforms and become text the recommendation and search systems can read. A music-only reel gives them nothing.
- Accessibility, via the auto-caption track that the transcription produces.
- Warmth without a face. A voice is a person. It is the closest a faceless account gets to the recognition a face provides, and it costs nothing to add.
- Nuance text cannot carry. Emphasis, doubt, humour — the things a caption flattens.
Write the script first
Improvising produces filler, and filler is what makes a voiceover sound amateur. Write it, read it, cut it, then record.
| Video length | Words | At ~2 words/sec |
|---|---|---|
| 10s | 18–22 | Tight |
| 20s | 36–44 | Comfortable |
| 30s | 55–65 | Room to breathe |
| 45s | 80–95 | Needs structure |
- Two words a second is a natural pace. Faster reads as rushed, slower as a lecture.
- Front-load the point. The first sentence is the hook, and it is competing with a thumb — reel hooks.
- Short sentences. Anything you cannot say in one breath is two sentences.
- Cut every hedge. "So basically", "I just wanted to", "kind of". They survive in writing and are unbearable aloud.
- Read it out loud before recording. Half the script will change, and that is the point of the pass.
- Say the searchable phrase once, plainly. It costs one sentence and it is the whole SEO benefit.
Recording it on a phone
Your phone's microphone is good. The room is the problem, and the room is free to fix.
- A hand's width from your mouth, slightly off to the side so plosives do not hit the mic directly.
- Soft room. A bedroom with curtains, a sofa, a rug. Kitchens and bathrooms are the worst places in the house — all hard surfaces, all echo.
- The classic trick works: record sitting in front of a wardrobe of hanging clothes, or with a duvet over your head and the phone. It sounds absurd and it is genuinely the difference between usable and roomy.
- Turn off everything that hums. Fans, extractors, air conditioning, the fridge. You will not hear it live and you will hear it in the file.
- Record two seconds of silence at the start, so a noise-reduction tool has a sample of the room.
- Do three takes and use the second. The first is stiff and the third is tired.
A cheap lavalier or USB microphone is a real improvement and is the last thing to buy, not the first. A phone in a soft room beats a good microphone in a kitchen every time.
Mixing it against the music
The most common failure in an otherwise good video: a voice sitting at the same level as the track, so neither is comfortable. The voice has to be clearly on top.
| Element | Relative level | Note |
|---|---|---|
| Voice | 0 dB — the reference | Peaks around -3 to -6 dB |
| Music under speech | -12 to -18 dB | Present, not competing |
| Music with no speech | -6 dB | Comes up in the gaps |
That last row is the technique worth knowing: drop the music while the voice is speaking and let it come back up between sentences. It is called ducking, most editors can do it automatically, and doing it by hand at three or four points in a thirty-second video is not much work.
Voiceover and beat-matching together
They coexist, with one adjustment: slow the cutting down. A cut every half second under someone speaking is exhausting, because the eye and the ear are being asked to track two different rhythms.
- Cut on every second or fourth beat when there is a voice, not every beat — how many cuts a reel should have.
- Let the sentences guide the shots. A new idea in the script is a good place for a change of subject on screen.
- Do not cut mid-word on a hard visual change; it draws attention to the edit.
- Keep the music's structure intact. A drop still lands well — use it under a sentence that deserves emphasis.
Synthetic voices
They are cheap, fast, and increasingly disliked. Audiences recognise the common text-to-speech voices immediately, and on a small account the recognition works against you: a synthetic voice reads as content produced at volume rather than by a person.
The honest position: your own imperfect voice beats a polished synthetic one for an account trying to build trust, and there is no version of a faceless brand that is harmed by an anonymous voice. If you genuinely cannot record — language, accessibility, a shared workspace — a synthetic voice is better than silence, and burned-in captions matter even more: captions on reels.
A ten-minute workflow
- Write 55 words for a 30-second video. Read it aloud, cut the hedges.
- Record three takes on a voice memo, in the softest room you have.
- Pick take two. Trim the silence at the ends, keep the two seconds of room tone.
- Drop it over the edit, duck the music to about -15 dB underneath it.
- Slow the cuts to every second beat.
- Watch it once on a phone, at low volume, the way most people will.
Then batch it. Recording six scripts in one sitting takes barely longer than recording one, because the setup and the warming-up are the expensive parts — which is the same logic as everything else here: batch-creating reels.