Post Production · Free Tool
“The music is too loud, I can’t hear you” is the note you get after the video is already up.
Drop in your voice stem and your music bed, exported at the levels you actually mixed them. This measures both to ITU-R BS.1770, reports the Loudness-Speech Ratio, and tells you how many dB to move the music. Nothing is uploaded — the decode and the maths happen in this tab.
🎙 Voice / dialogue stem
Speech only, no music. Drop a file or click.
🎵 Music bed
Music and effects, no speech. Drop a file or click.
How loud should background music be under a voiceover?
Aim for a Loudness-Speech Ratio of 5 LU or less — the difference between the loudness of the whole programme and the loudness of the speech in it. Under anchor-based loudness normalisation that is the guidance for most productions: it keeps dialogue intelligible while still letting music and effects be louder where speech is not carrying the scene. Two more published numbers worth knowing: Netflix asks for a programme loudness range of 4–18 LU but 7 LU or less for dialogue, and YouTube dialogue-led content generally sits around −15 to −13 LUFS integrated.
Why “put the music 20 dB under the voice” is the wrong question. A fader reading tells you what you did to a signal, not what a listener receives. A sparse piano bed and a dense compressed drum loop at the same fader setting are nowhere near the same perceived loudness, and the busier one is the one eating your consonants. Loudness measurement exists precisely because peak and fader values do not predict what people hear. There is no single correct fader number, because it depends on the material. There is a target for the relationship between speech and programme — that is what this measures.
What this does and does not claim
Three published ones, and nothing else. LSR ≤ 5 LU is the guidance for most productions under anchor-based loudness normalisation. Netflix asks for a programme loudness range of 4–18 LU but ≤ 7 LU for dialogue. YouTube dialogue-led content sits around −15 to −13 LUFS integrated. Each verdict below names which one it is judging against.
Because separating speech back out of a finished mix is an estimate, and every number after that inherits the error. You still have the stems in your editor — measuring them directly is exact. Export both at the levels you actually mixed them; the programme figure is their sum.
No. The file is decoded by your own browser and the analysis runs in this tab. Nothing is sent to a server, and there is nothing here to send.
The loudness figures are a real BS.1770 implementation — K-weighting, 400 ms blocks at 75% overlap, the −70 LUFS absolute gate and the −10 LU relative gate. Range uses EBU Tech 3342: 3-second blocks, −20 LU gate, 10th to 95th percentile. Audio is resampled to 48 kHz first, because the filter coefficients are specified there.