Guide

How to mix music under dialogue so both survive

Turning the music down is the fix everyone reaches for, and it is the worst of the three. The problem is not level, it is that speech and music are competing for the same narrow band, and there are better ways to give speech that band back.

For anyone cutting dialogue: film, documentary, corporate, podcast, YouTube.

Where the collision actually happens

Speech intelligibility lives in a surprisingly narrow band. The vowels that carry loudness sit low, but the consonants that carry meaning, the t and k and s and f sounds, live roughly between 1 kHz and 4 kHz. Lose that region and speech does not get quieter, it gets unintelligible, which is a different and worse problem.

Almost every piece of music has a great deal going on in exactly that band: guitars, snare, synth leads, and worst of all, other people singing.

The low end is not the issue. A bass line at 60 Hz masks nothing in a voice. This matters, because the standard reflex of high-passing the music to make room is aimed at the wrong end of the spectrum. You lose the weight and gain no clarity.

The single biggest lever

Choose music with nothing in the middle. A pad, a bass and a texture will sit under dialogue at a healthy level and never fight it. A track with a vocal will fight at any level, because the brain insists on trying to parse both voices at once. No amount of mixing rescues a vocal track under dialogue.

The three fixes, best first

1. Arrange around it

Free, invisible, and by far the most effective. Cut the music so its busy sections land in the gaps between lines, and its quiet sections land under speech. Start a music cue on the last word of a line rather than the first. If a track has a section without the midrange instruments, loop that section under the dialogue and save the full arrangement for the wide shot with nobody talking.

This is editing, not mixing, and it is why a well cut scene needs almost no processing.

2. Carve a hole

A broad, gentle dip in the music where the consonants live. Start with a bell around 2 kHz, about 3 dB down, with a wide Q, and adjust by ear. Wide and shallow is the goal: a narrow deep notch is audible as a hole in the music, whereas a broad shallow one just sounds like the music sitting slightly further back.

Do it on the music, never on the voice. And if you find yourself needing more than about 5 dB, the real answer is fix 1 or a different track.

3. Duck it, slowly

A compressor on the music, keyed from the dialogue. Everybody knows this one and most people set it wrong.

Ducking is the last resort of the three because it is the only one the audience can hear working. Used gently on top of the other two, it is invisible.

Levels to aim at

Two separate questions: how loud is the programme overall, and how far below the dialogue does the music sit.

Programme loudness

Where it is goingTargetStandard
European broadcast-23 LUFS integratedEBU R128
US broadcast-24 LKFS integratedATSC A/85
YouTube, Spotify and most streamingaround -14 LUFS Platform normalisation, not a delivery spec
Podcast-16 LUFS stereo, -19 LUFS monoCommon practice

Note the size of that gap. A mix delivered at -23 for broadcast and then uploaded to YouTube gets turned up by about 9 dB, and every masking problem you got away with becomes audible. Check the mix at the level it will actually be heard.

Music against dialogue

There is no standard, only a working range. When the words matter, music typically sits 12 to 18 dB below the dialogue. Under a montage with no speech it comes up to within a few dB or takes over completely. The move between those two states is what makes a scene feel scored rather than soundtracked.

The only test that matters

Play it on a phone speaker at low volume. Phone speakers have almost no bass, so the midrange fight is all that is left, and any masking problem becomes obvious immediately. If the dialogue survives there, it survives anywhere.

When the music is meant to be in the scene

Everything above assumes score, sitting outside the world of the film. Source music, playing from a radio or a room down the hall, is a different job: it needs to be placed rather than balanced, and once it is properly placed the masking problem often solves itself, because a wall has already removed the midrange for you.

Two related guides

If the music in your scene is coming from somewhere in the world, how to make music sound like it is playing in another room covers the filtering, and what counts as source music covers when to treat it that way at all.

THALM does that placement as geometry rather than as EQ, and it follows your playhead, so the treatment moves with the shot instead of being keyframed.

More guides like this

We write these when there is something worth writing down. One email when a new one lands or a new Tunary instrument ships. No newsletter, no schedule.

See all guides