Short answer: a de-esser is a frequency-specific compressor that turns down harsh sibilance, the "sss" and "shh" sounds, usually somewhere around 4 to 10 kHz, without dulling the rest of the voice. It listens to that one band and only acts when the sibilance spikes, so the vocal keeps its brightness and clarity everywhere else and only the sting comes down.
The exact band, the threshold and the amount all depend on the voice, the mic and how bright the rest of your chain is, so treat the frequency figures here as starting zones and set them by ear. None of this is tied to one product.
A de-esser is a frequency-specific compressor that turns down harsh sibilance, the "sss" and "shh" sounds in a vocal, usually somewhere in the 4 to 10 kHz range, without dulling the whole voice. It is aimed at one narrow problem: the moments when an "s" or a "sh" jumps out sharp and piercing, while it leaves the rest of the word alone.
The reason you reach for a de-esser instead of a plain EQ cut is that sibilance comes and goes. It only appears on certain consonants, for a fraction of a second at a time. A fixed EQ cut in that range would dull every ess and pull the brightness out of the whole vocal, all the time, including the parts that were never harsh. A de-esser turns the band down only in the instant the sibilance spikes, then lets go, so the voice keeps its air and clarity everywhere else.
Under the hood it works like a compressor with its ear tuned to one band. It watches the level of the chosen high-frequency band against a threshold you set. While that band stays under the threshold nothing happens. The moment a sharp "s" pushes the band over the threshold, the de-esser turns it down for that brief moment, then releases as the sibilance passes. That is why it only acts on the essers and not on the vowels or the body of the voice.
A de-esser is not a special effect. It is an automatic, very fast volume dip that only listens to the harsh band. Everything it is good for comes from that one idea: watch a narrow slice of the top end, and pull it down for the split second it turns harsh.
Sibilance is the burst of bright, hissy high-frequency energy that certain consonants make, mainly the "s", "sh", "z", "ch" and "t" sounds. It is a normal and necessary part of speech, the thing that makes words crisp and easy to understand, and you do not want to remove it. You only want to tame it when it turns sharp enough to sting.
That energy usually sits somewhere in the 4 to 10 kHz range, and it moves around with the voice and the consonant. A hard "s" tends to sit higher, often around 6 to 9 kHz, while a softer "sh" or "ch" usually sits a little lower. Deeper voices tend to sibilate lower and brighter voices higher, so there is no single frequency that fits every singer.
Sibilance turns harsh for a few common reasons that stack up. Close mic positions and bright condenser mics emphasise the top end. Compression turns the loud body of the voice down and, with makeup gain, effectively lifts the quiet high-frequency detail, so the essers come up with it. And the presence and air lifts people add with EQ to help a vocal cut through also raise the exact band sibilance lives in. Do all three, which is normal on a modern vocal, and a take that sounded fine raw can start spitting.
Because the steps above make sibilance worse, a de-esser usually sits after the compressor and the bright EQ in the vocal chain, cleaning up the sibilance those steps raised. Some engineers de-ess earlier, before the compressor, so it does not clamp hard on the loud essers in the first place. Both are common, and the right spot is wherever the sibilance is actually being made worse.
Find the frequency where the harshness lives, set the threshold so the de-esser only acts on the harsh peaks and not on every word, and use the least reduction that clears the problem so the voice does not start to lisp. Three moves: aim it, set how easily it triggers, then set how hard it pulls.
| Control | What it does | A place to start |
|---|---|---|
| Frequency / band | The slice of the top end the de-esser listens to and turns down. You sweep it to sit right on the harshest sibilance. | Sweep around 5 to 9 kHz to find the sting |
| Threshold | The level the sibilant band has to cross before the de-esser acts. Set it so normal words pass and only the harsh essers trigger it. | Lower until only the essers show reduction |
| Range / amount | How far it turns the band down once it triggers. Too much and the voice starts to lisp. | A few dB, the least that clears it |
| Mode | Split-band pulls down only the chosen band. Wideband dips the whole signal for the instant the band spikes. | Split-band for the most transparent result |
| Listen / solo | A monitor mode that plays only the band being reduced, so you can aim the frequency by ear. | Use it to aim, then switch it off |
Sweep to the harshness. Start by finding the band. Most de-essers have a listen or solo mode that lets you hear only the part being reduced. Turn it on, play a line with strong essers, and sweep the frequency until the "s" sounds are loudest and most obvious. That is where the sibilance is strongest, so that is where you aim. Then switch the listen mode off.
Set the threshold. Lower it until the de-esser starts turning down the harsh essers, and watch the reduction meter: it should blink only on the "s" and "sh" moments and stay still on the rest of the words. If it is clamping on vowels or on every syllable, the threshold is too low and you are de-essing the whole vocal, not just the sibilance.
Use the least that works. Finally, set how much it pulls down. A few dB on the worst peaks is usually enough. Push it too far and you strip the "s" out of the voice entirely, so the singer sounds like they have a lisp or a cold. Good de-essing is a subtraction you can barely hear: the harshness gone, the words still crisp. When in doubt, back it off and compare against the bypassed vocal at a matched level.
Split-band or wideband. Most de-essers offer two modes. Split-band de-essing only turns down the chosen high band and leaves the rest of the voice untouched, which is the more transparent choice and the usual default. Wideband de-essing briefly turns the whole signal down when the band spikes, like a fast compressor keyed to the essers. Wideband can sound more natural on some voices because it keeps the tonal balance intact for that instant, but it can also cause tiny level dips if you push it, so split-band is the safer place to start.
You do not strictly need a dedicated de-esser. You can tame sibilance by hand with clip gain, ride it out with volume automation, or use a dynamic EQ to dip the harsh band only when it spikes. Each one gets to the same place: turn the sibilance down for the moments it is harsh, and leave the rest of the vocal alone.
Clip gain by hand. The most surgical option is to do it manually. Zoom into the vocal, find each harsh "s" in the waveform, and pull just that little slice down by a dB or a few with clip gain or a volume automation node. It is slower, but it is the most transparent method there is, because you are only touching the exact moments that are harsh and nothing else. For a vocal with only a handful of standout essers, this is often the cleanest fix.
Dynamic EQ. The closest automatic alternative is a dynamic EQ. Set a band over the sibilant range, roughly 4 to 10 kHz, and tell it to dip only when that range gets loud. That is essentially what a de-esser does, so a dynamic EQ makes a very good stand-in, with the bonus that you can aim the band precisely and even use more than one, a higher band for a hard "s" and a lower one for a "sh".
Blunter options. A plain, fixed EQ cut in the sibilant range works in a pinch, but it dulls every ess and the general brightness all the time, not just the harsh peaks, so it is a blunt fix rather than a clean one. A multiband compressor set with one band over the sibilant range is another route, and is really the same idea as a de-esser under a different name. Whichever you use, the principle does not change: act only on the harsh band, only when it spikes, and only as much as it takes.
A de-esser is a scalpel, not a repair for a bad recording. If a vocal is spitting badly, some of it may be the mic, the distance or the room, and moving the singer slightly off to the side of the mic or backing off it a touch can cut sibilance at the source. Fix what you can before the take, and let the de-esser handle the little that is left.
A de-esser works on the voice itself, but it has an easier job when nothing else in the mix is crowding the same bright band. If the backing is piling energy up around 4 to 10 kHz, the vocal has to fight for the top end and every harsh ess stands out more. The calmer and more focused the arrangement is up there, the less there is to fight in the first place.
We do not make a de-esser, so this is not that. What we make is instruments. The Collection is our three of them together, ARGISH, SILT and REHEAT, with three separate licence keys. They make clean, focused backing sounds that stay in their own part of the range instead of piling up in the 4 to 10 kHz band where sibilance lives, so the voice sits on a clean bed and its harsh moments are easier to tame. ARGISH is a self-playing drone synth, SILT is a tape-loop instrument, and REHEAT writes acid lines. Every one ships as AU, VST3, AAX and a standalone app on macOS, signed and notarized, and a VST3 on Windows.
You can hear one free in your browser first, no install and no account, then decide.
We write these when there is something worth writing down. One email when a new one lands or a new Tunary instrument ships. No newsletter, no schedule.