Back to index
Sheet 01 / Field Notes

Acoustic Echo Cancellation

// On the curious craft of removing your own speakers from your own microphone - a tour through WebRTC's AEC3.

I don't think most people have heard of this, or I at least never thought about it before I had to build out a somewhat proper VOIP pipeline but audio coming from your computer without any kind of processing would interfere with the audio coming from the speakers (duh).

How does that work?

COMPUTER[ AEC3 ]render x[n]SPEAKERecho path / ~ 10-100 msMICNEAR-ENDyour voice v[n]capture y[n] = v[n] + (h ∗ x)[n] + noise
Fig. 1 - The loop AEC has to undo D-1.1

Your microphone is not inherently able to sample only your voice. Any loud enough sound in the room can be sampled too, including audio from the speaker. This is why microphones and operating systems often include built-in filtering (either from the OS itself or through vendor drivers/software) to handle this. The algorithm conceptually works by understanding that the signals between the microphone (capture path) and the speakers (render path) become mixed, and it is possible to "remove" the render path signal from the capture path (through some echo cancelling model). In fact, there is open source software called WebRTC which has an audio processing module called AEC3 that does this exact thing, even though on a much more advanced level. Let's try to explain its architecture and most important modules.

The whole problem fits on essentially one line. Call the far-end signal we're about to play out the render , our own voice the near-end , and what the microphone actually sees . Then

(0.1)

where is the unknown impulse response of the speaker -> room -> mic path and is background noise. The job of an echo canceller is to estimate well enough to subtract out of and hand back just . Seems easy but the tricky part is the vast amount of unknowns: delay, shape of the impulse, how much is actually echo vs near-end audio (your own voice).

§ 01 Finding the delay

First we assume that the render audio usually comes before its echo appears in the capture stream. The exact delay depends on hardware buffering, OS scheduling, and the acoustic path, so AEC3 has to estimate it adaptively instead of treating it as a fixed number. The source does this with a render delay buffer, a decimated copy of the signals, a bank of matched filters, and a histogram-style lag aggregator.

The loop, roughly

  • Render arrives from the application and gets stored in a ring buffer.
  • It is decimated down to a lower sample rate. The config contract accepts a downsampling factor of 4 or 8, defaulting to 4, and quietly rewrites anything else to 4. At 48 kHz with factor 4 the delay search runs against a 4 kHz copy of the render, which is more than enough resolution to localise an echo and a lot cheaper to run a filter over.
  • Capture arrives, gets decimated the same way, and is handed to a bank of short adaptive filters - the matched filters, five of them on default config.
  • Each filter is scored independently: it runs NLMS against the render at a different alignment, then asks "how much of the capture's energy did I just remove?" The filter that removed the most wins for this block.
  • A lag aggregator watches all the per-filter winners over time and only commits to one when the same lag keeps coming up.
  • If the committed lag stops moving for more than 1/2 second, the matched filters are zeroed and restarted. This is a separate thing from the estimate being marked trustworthy, which happens on vote count alone (more info on both later below).

Three bands, and only one of them gets a canceller

Modern capture devices run at 48 kHz, but AEC3's core block processing runs at a 16 kHz band rate. At 48 kHz the full-band audio is split into three 16 kHz bands:

All three bands are carried through the same framing machinery, and the pipeline operates at the band rate in fixed blocks:

A block is samples = 4 ms at the band rate, which is also the half-FFT length the canceller uses downstream. Granted here is some interesting behaviour which in fact makes sense: the echo model only ever sees band 0. There is exactly one adaptive filter pair per capture channel and it runs on 0 to 8 kHz. Nothing above 8 kHz is ever convolved, correlated, or predicted (naturally, due to human speech being more or less restrained to 8kHz).

The upper two bands do still get processed, just very cheaply, in that they are handed one scalar gain per block which gets applied flat across all 64 samples. That gain moves around, since it collapses to 0.001 when the echo saturates and an anti-howling term drags it down to whenever the high-band render energy climbs above the low band's, but the term that dominates in normal operation is just the minimum per-bin gain that the suppressor already picked for the 4 to 8 kHz half of band 0. (A fourth term clamps it whenever low-frequency echo is present, though its default value is 1.0, so it sits there doing nothing until you tune it.) The top two thirds of your spectrum therefore get attenuated by inference: whatever band 0 concluded about its own upper half, the high bands inherit.

The delay search from the bullets above lives on band 0 as well, mixing or selecting channels down to a single alignment stream and then decimating again by the configured factor, so with the default factor of 4 the matched filter is hunting for an echo on a 4 kHz signal, and factor 8 would make it 2 kHz.

ROW 1 / FINDING THE DELAYRENDER x10 ms frame, cut into 64-smp blocks3-bandsplitband 0ring bufferread ptrFFT 128/44 kHzDELAY SEARCH5 NLMS matched filters250-block sliding voteopened up in Fig. 3committed delay aligns the read pointerROW 2 / CANCELLING ITCAPTURE y3-bandsplitband 0, y[n] — 64 samples per blockcapture, decimatedX(k, ω)MAIN 13 + SHADOW 13 PARTITIONS832 taps = 52 ms of roomoverlap-save: keep the last 64 of 128ΣŶ(k, ω)IFFT 128keep the last 64ŷ[n]+−e[n]FFT 128Hann + zero-padadapt: H ← H + X* · G (2.3)G, per (2.5)LINEAR OUTPUT (opt.)e[n] as-is, band 0, 4 msFFT 128√Hann on [e(k−1); e(k)]gain G, 65 bins in [0, 1]ω = 0ω = 64residual echo + reverbλ = 0.83 fixedshape learnedshaped noise N × √(1−G²)IFFT + overlap-add√Hann, 50% hopbands 1 and 2 bypass all of row 2g, one scalar per blockmostly min G over 4 to 8 kHz3-bandmergeOUTPUT 48 kHznear-end intact
ctrl + scroll, or double-click, to zoom 100%
Fig. 2 - Signal path on default config. The linear filter and the 65-bin gain live on band 0. Note where the frequency domain ends: the partition sum is inverse-transformed first, the subtraction happens on samples, and each consumer of e[n] builds its own forward transform. The suppressor's input is the raw capture spectrum instead whenever the state machine distrusts the linear filter. D-1.2

The matched-filter bank

The name is a little misleading, because a matched filter in the textbook sense correlates a known template against a signal, whereas the "simple echo cancel algorithm" inside each of these is a short NLMS adaptive FIR, the same family as the real canceller in §02 and just tiny and running on decimated audio. Each filter covers a 32-sub-block window into the decimated render buffer, and consecutive filters are shifted by 3/4 of that window (24 sub-blocks). So if you imagine a long render history to the left and "now" on the right, the bank is a row of overlapping windows, each one asking "if the echo lived inside my window, how well could I cancel it?".

Concretely, for filter sitting at alignment shift , with the window of decimated render it currently sees and the decimated capture, every sample of the sub-block does one NLMS step:

(1.1)

with by default, and the step skipped entirely when the render window is too quiet to excite the filter or the capture sample is clipping. The whole block then gets scored, and the lag falls straight out of the filter taps themselves:

(1.2)

is the energy the filter actually removed, measured against the capture block's own energy as the anchor. The per-block winner is the largest , which is the same as saying the smallest residual. A filter is only allowed to compete if it is marked reliable, which needs two things: its residual has to sit below 20% of the anchor, and its peak has to be off the boundary of its window, since a peak pinned to either edge usually means the true lag is outside this filter's window and the argmax is an artefact.

The reason any of this works is that an NLMS filter driven to convergence on a broadband reference approaches the normalised cross-correlation, so the argmax of does land on the lag you want even though the code never forms a correlation anywhere along the way. That difference leaks into the details if you go reading the source, since the quantity being scored is residual energy, which is why reliability comes down to an energy threshold and the height of the peak never enters into it at all.

// Sketch of the matched-filter scoring loop, once per block.fn estimate_delay(render: &DownsampledRender, capture: &[f32]) -> Option<Lag> {    let anchor = energy(capture); // the sum of y^2 everything is measured against    for (i, h) in filters.iter_mut().enumerate() {        // The decimated render is written into the ring buffer backwards, so        // each filter starts at its own shift and walks it in reverse.        let start = (render.read + shift[i] + capture.len() - 1) % render.len();        let error_sum = nlms_run(start, render, capture, h);        let peak = argmax_abs(h);        lag_estimates[i] = LagEstimate {            lag:      peak + shift[i],            accuracy: anchor - error_sum,        // energy removed, not correlation            reliable: error_sum < 0.2 * anchor   // residual well below the anchor                      && peak > 2 && peak + 10 < h.len(), // peak off the boundary        };    }    // Winner is the highest accuracy among the reliable ones. Then it votes.    aggregator.aggregate(&lag_estimates)}

The aggregator on top of this is what actually commits to a delay. A one-block lag estimate is cheap and noisy; what you want is a lag that keeps showing up. So the winner of each block is dropped into a sliding window of the last 250 votes, and the aggregator reports whichever lag currently holds the most votes. Crossing 5 votes buys you an initial, provisional estimate, and crossing 20 gets it marked refined, which is a latch, because once any lag has been refined the provisional path switches off and nothing under 20 votes is ever emitted again.

Two things here are very easy to run together, and keeping them apart matters. Being refined is purely about vote count, and its consumer is the delay controller, which only applies its hysteresis when both the previous and current estimates are refined. The half-second rule is a different mechanism with a different counter: if the committed lag comes back identical 126 blocks in a row, at 250 blocks/s, the estimator zeroes the matched filters and starts them over. It does not check quality, so a provisional lag that happens to repeat will trigger it just the same. The nice detail is what survives that reset: the filters are wiped but the vote histogram is not, so the estimator forgets how it found the delay while remembering that it did. There is also a clock-drift detector running on the side, fed only by refined estimates, because on real hardware the render and capture clocks rarely tick at exactly the same rate.

decimated render, 4 kHzdecimated capture = render delayed by τ, plus noise5 matched filters, each 96 taps, shifted 72 taps apart. Curve = the filter's own taps h.f0unreliablef1unreliablef2unreliablef3unreliablef4unreliableenergy removedτ = 150votes held in the last 250 blocks  ·  bar height on a √ scale20 refined5 provisionallag 0lag 384 = 4 shifts + one window
ctrl + scroll, or double-click, to zoom 100%
block0
winner —
committed—
quality none
held 0/126
filter resets0
Fig. 3 - The filter bank and the vote, running live D-1.3

That figure is running the actual algorithm live in your browser, so it repays reading in layers. The five lanes in the middle are the bank, drawn on a shared lag axis so you can see the windows overlapping by a quarter of their length, and the wiggly curve inside each lane is that filter's own tap vector as NLMS pushes it around. The dot on each lane is that filter's , and the shaded strips at the two ends of every window are the guard bands, so a dot landing inside one of those is what gets a filter marked unreliable. Bars on the right are the scores , the capture energy each filter managed to remove.

Drag and watch the order things happen in. One lane starts winning almost immediately, but the committed lag at the bottom stays empty until a single bin has accumulated more than 5 votes, and only crosses into refined past 20. Then leave it running and the interesting thing shows up on its own: every so often the tap curves all collapse flat and rebuild themselves, which is the half-second consistency rule firing, zeroing the filters while the vote histogram underneath keeps every one of its votes. The reset counter in the readout ticks up each time. Tap counts are scaled down for legibility, but the window-to-shift ratio, the bank size, the 20% reliability threshold and the 5 and 20 vote thresholds are all the real numbers.

Aside: what if you have more than one channel?

Everything above is drawn mono-in, mono-out for the sake of clarity. Real AEC3 doesn't insist on mono, i.e it accepts any number of render channels and any number of capture channels and sizes itself accordingly at construction. (Only at construction, mind you - there is no reconfiguration API, so changing channel count means building a new instance.) Two interesting things happen as you scale up:

On the delay search, the matched-filter bank does not run once per channel. Both sides first pass through an alignment mixer that produces one stream for delay estimation. Depending on configuration, it can average channels, use a fixed first channel, or adaptively select the strongest alignment channel. The default configuration is to use adaptive selection rather than literal downmixing as searching independently per channel would mostly find the same device delay several times and let the estimates disagree.

On the main + shadow filters from §02, the picture fans out. AEC3 instantiates one pair of adaptive filters per capture channel, and each of those filters carries FFT-domain partitions for every render channel. The per-block echo prediction for capture channel is the sum over render channels of the channel-pair's convolution:

(1.3)

So with two render channels and three capture channels you end up with three main filters and three shadow filters in the subtractor, each holding two render-channels-worth of FFT partitions. That is six channel-pair convolutions per filter bank, and since the main and shadow banks both run every single block, twelve convolutions per block in total. The update rule is the same one you'll see in §02, just applied per channel-pair.

With the above said, for the rest of this post I'll stay in the mono-mono case, since the math is identical with just fewer subscripts.

§ 02 Actually cancelling the echo

So now we actually know the echo path delay and we can move onto trying to actually echo cancel on a more qualitative basis with the main filter heavy lifting.

Why everything runs in the frequency domain

A room is a long, ugly impulse response. A real loudspeaker -> mic path can easily be 30 to 100 ms of reverberant tail, which at 16 kHz is 480 to 1600 samples. A time-domain FIR of that length, run sample by sample, is quite expensive, so the whole thing moves into the frequency domain where a convolution collapses into a pointwise multiply.

Moving it there costs you something first, and the fix has a name that matters here, because §03 solves the same problem a different way. An FFT has no concept of "before the block" or "after the block"; it treats the samples you hand it as one period of something that repeats forever. Multiply two spectra, transform back, and what you get is the circular convolution, where the tail that should have spilled past the end of the block instead wraps around and lands on top of the beginning. The first stretch of every block comes out contaminated.

The trick called overlap-save handles this by being deliberately wasteful. Hand the transform 128 samples at a time, made of the previous 64-sample block followed by the current one, so successive frames overlap by half. Multiply, transform back, then discard the first 64 samples and keep only the last 64. The wraparound can only ever reach a distance equal to the impulse response length, and each chunk of the filter is exactly 64 taps long, so all of the damage is confined to the half you just threw away. What is left is the true linear convolution, sample for sample. You save the good half and bin the rest, which is where the name comes from, and note that nothing is faded or added together anywhere: the output is either kept whole or dropped whole.

That still leaves the length problem, since a 128-point transform is nowhere near long enough to hold a whole room response. So the impulse response is sliced into -sample chunks called partitions, each transformed once, and the echo prediction for this block becomes the sum of every partition multiplied by the render block it lines up with:

(2.1)

In this notation, is the block index: the first 4 ms block, the next 4 ms block, and so on. means "one frequency bin of the FFT". So is not one sample of audio; it is the render block after the FFT, viewed at one frequency. In practice is a discrete bin index running over 65 bins, but writing keeps the equation readable. The index counts backwards through the partitions, so multiplies the newest render block and multiplies one from twelve blocks ago, and adding them all up is the convolution.

in the default config, which at 64 samples a partition means the linear filter spans 832 samples, or about 52 ms, comfortably shorter than the 100 ms tail quoted above. So the linear filter is only ever trying to cover the early part of the room response, and cleaning up whatever it cannot reach is the job of §03.

Each block, AEC3 does one FFT of the new render frame, multiplies it through the partitions, sums, inverse-transforms, and gets a prediction of the echo. Subtract from and you get an error signal:

(2.2)

That error is what the canceller emits as its linear output, meaning capture with the predicted echo torn out of it.

Lets be pedantic for a second about where that subtraction actually happens, because the notation in (2.2) can mislead. The prediction gets inverse-transformed first, and the subtraction is then done in the time domain, one sample at a time, against the captured samples. So what leaves this stage is not a spectrum at all: it is 64 real samples per block, a signal you could write to a wav file and listen to. (2.2) is the same statement viewed through the transform, which is fair because the transform is linear, and it is easier to read next to (2.1). Just do not picture a spectrum being carried forward, because the next two stages each build their own.

That also settles what you get if you ask the canceller to hand you the linear output directly, which it can do. You get precisely those samples: the raw time-domain error, with whichever of the two filters was selected and the short crossfade between them, clamped to 16-bit range and otherwise untouched. Every window mentioned anywhere in this post is applied to a copy. Both the adaptation's transform and the suppressor's take the error by reference and write into buffers of their own, so neither one ever modifies the signal itself, and the exported linear output carries no window at all.

The main filter and the shadow filter

As explained before every capture channel actually has two adaptive filters running side by side:

  • The main filter is the steady-state path. In the default config it has 13 FFT-domain partitions after the initial phase, and its update gain is conservative so a single weird block does not knock it off course.
  • The shadow filter is the faster challenger. In this port it is not shorter by default as it also uses 13 partitions after startup - but its update rate is much more aggressive. When the echo path changes, the shadow can become the better residual output before the main filter catches up.

Each block, AEC3 computes both residuals. If shadow-output usage is enabled, which it is by default, the output selector can choose the shadow residual when it is clearly better than the main one, with a short crossfade between choices. In the subtractor itself, the shadow is also kept practical, i.e if it is worse than the main for five blocks running, it is copied from the main and given a reset hangover.

Both filters share the same update shape, in that each block produces a gain spectrum and every partition then gets nudged by it against the render block it lines up with:

(2.3)

The conjugate is the whole trick, since multiplying by is correlation in the frequency domain, so each partition moves in the direction that explains whichever part of the error its own slice of render history is responsible for. Lets focus a bit again on the peculiarities of the fourier domain and complex numbers: The filter adapts every single block, and nothing in the update rule by itself keeps a partition's impulse response inside its first 64 taps. I.e this means it can creep longer and the wraparound starts reaching into the half you were about to keep, quietly poisoning the linear output. So after each update the canceller takes one partition back into the time domain, erases every tap past the sixty-fourth, and transforms it forward again, working through the partitions in rotation so each one gets cleaned every thirteen blocks (i.e the drift will be small enough so that it wont matter). That small piece of housekeeping is what keeps the discard-the-first-half shortcut correct while the filter is still moving. Where the two filters diverge is entirely in how gets computed, and the two computations have almost nothing in common beyond that shared shape. The shadow filter is textbook NLMS:

(2.4)

A fixed step size over the render power summed across partitions, zeroed outright in any bin where that power sits below a noise gate, and that gate is doing the job normally does in the textbook version, just as a hard cutoff on the whole update. The main filter:

(2.5)

There is no fixed learning rate here at all, because is a per-bin running estimate of how badly the filter is currently misaligned and it takes over the role plays above, so confident bins take small steps while bins that look wrong take large ones, which puts this much closer to a Kalman gain than an NLMS one. It gets decremented by each update it authorises, then grown again by a leakage term, then clamped into .

The in (2.5) is a bit different than (2.2) asa it is a transform/edited signal the subtractor builds for itself, out of the same time-domain , and it is built differently: the 64 error samples get a Hanning taper and are dropped into the upper half of a 128-point buffer with the lower half left as zeros. The zero padding is the overlap-save idea again, in a different costume. The gradient formed from this is a correlation of the error against the render history, and padding to twice the length is what keeps that correlation linear instead of letting it wrap around the frame. The Hanning taper is there because the same spectrum doubles as the per-bin power estimate in the denominator of (2.5), and a taper keeps energy from one bin leaking into its neighbours and corrupting that estimate.

So the same 64 error samples get transformed twice, with different windows, for different consumers, and there is a clean reason the windows differ. This one produces a coefficient update, and a coefficient update never has to be turned back into audio, so no reconstruction property is required of it and the window can be chosen purely to make the estimate good. §03's transform of the very same produces the audio you actually hear, so it is stuck with a window that reconstructs exactly, which turns out to be a real constraint.

The shadow filter is additionally used to moulate the leakage term discssued ealrier, as it is chosen per bin by comparing the two filters: in any bin where the shadow's residual beats the main's, leakage jumps by a factor of a thousand, from 0.00005 to 0.05. So the main filter's misalignment estimate inflates exactly where the shadow is outperforming it, which inflates , which makes the main filter re-adapt aggressively in precisely those bins.

As such the shadow filter acts as the main filter's divergence detector, running continuously and per frequency bin. Also that explains the reset hangover from a paragraph ago, because right after the shadow has been copied from the main it would trivially tie with it and start poisoning the main's leakage decisions, so the comparison gets suppressed for 25 blocks while the shadow re-adapts on its own. What you end up with is a canceller that stays stable while things are stationary and still has something fast-moving watching every bin for the moment they aren't.

§ 03 The finishing touches

While the linear filter is what does most of the job, what makes a call sound "clean" is the residual-suppression stage that runs after it.

One number per bin, rebuilt every block

As said before, what arrives here is the time-domain (no Fourier domain stuff) , 64 real samples per block, and this stage builds the third transform of that same signal. The subtractor's own one, the Hanning-tapered padded spectrum behind (2.5), is no use for this, since it was shaped for estimating a gradient and carries no reconstruction property whatsoever. The suppressor has to hand back audio, so it starts over with its own forward transform, and that transform is the windowed one everything below turns on.

With that established, the suppressor is doing something very simple underneath all its estimators. It looks at the 65 bins of its transform and works out two numbers for each: roughly how much echo it thinks is still in there, and roughly how much near-end. The ratio between them becomes a gain somewhere between zero and one. Bins that look mostly like leftover echo get a small gain, bins that look mostly like your voice get something close to one, and the suppressed output is that gain multiplied straight in, bin by bin:

(3.1)

The window sits right there in the definition, applied to the 128-sample frame before the transform runs. That is the sense in which carries a factor of inside it, and it will matter again shortly. Everything in the rest of this section exists to make that pair of estimates less stupid than it sounds. But (3.1) also settles a question left open in §02, which is why the transform around this stage has to be arranged differently from the one around the linear filter.

Why this stage cannot use overlap-save

First, be precise about which two operations are being compared, because the interesting one in §02 is easy to miss. (2.2) is a subtraction and has nothing to do with any of this. The operation that matters there is buried one level down, inside the sum in (2.1): each partition contributes a spectrum multiplied by a spectrum, and there are thirteen such products added together. So the pair to hold side by side is

Both are a spectrum times a spectrum, so the multiplying is not what separates them. What separates them is the multiplier: in one case and in the other, and specifically what each of those looks like once you transform it back into the time domain. Call those time-domain versions and .

To reiterate a bit, the reason it matters is that a pointwise product of two length- spectra is never a plain convolution in Fourier domain, It is a circular one! In more detail:

(3.2)

The is the whole problem. Samples that should have run off the end of the frame reappear at the front instead. There is exactly one condition under which you may ignore that, and it is a condition on the multiplier's impulse response : if dies out within samples, then everything from sample onwards is untouched by the wrap and agrees exactly with the linear convolution.

(3.3)

Cleraly as seen, this is just overlap-save algo. Put and the filter satisfies it by construction, because a partition is exactly one 64-tap slice of the room response and is zero everywhere after that. So against , samples 63 through 127 are exact, and keeping the last 64 is provably safe. The tap-erasing housekeeping from §02 is nothing more elaborate than this condition being re-enforced by hand every block so it keeps holding while the filter moves.

But with its not so simple, for a reason visible in (3.1) itself: is a real number per bin. It is a gain with no phase attached, meaning a spectrum that is purely real has a symmetric inverse transform:

(3.4)

So is a zero-phase response centred on , which in a circular buffer of length means it has large values at and equally large values at , wrapped around to the far end of the frame. It occupies both ends at once. There is no short of for which it vanishes, (3.3) is therefore unavailable at any , every output sample carries some wrap, and no clean half exists to save (I like to think of it as one half of a "spike" in 0 and its reflection running back at 127). You could force short by erasing its tail the way the filter does, but that would turn into some other gain, which rather defeats the point of having computed it.

There is a second, independent problem on top of the first. is rebuilt from scratch every four milliseconds, so even if some clean region did exist, the region from one block and the region from the next were produced by different filters and would not meet cleanly at the join.

So it tapers instead

Since the wrap cannot be isolated, it gets made inaudible instead, and the window gets applied twice at two separate moments. The first is inside (3.1), on the way into the forward transform, which is why and everything derived from it already carry one factor of . The second comes after the inverse transform, at the moment the finished samples are handed out. The wrapped junk and the block-to-block mismatch both live at the frame edges, and a window that falls to zero at the edges attenuates them into nothing.

A frame is samples, which at the 16 kHz band rate is 8 ms, and it is made of the previous 64-sample block followed by the current one. A block is 64 samples, 4 ms, and that is how far the whole thing advances each time. So consecutive frames overlap by half, and each block only ever emits 64 finished samples even though it just inverse-transformed 128 of them.

Each block runs exactly one inverse transform, and what comes back is 128 samples long because that is the length of the frame that went in. Cut it down the middle and give the two halves names:

The head is used straight away. The tail is set aside and waits one block, and it is the only piece of state this stage carries between blocks. So output sample is last block's tail, still falling, plus this block's head, rising:

(3.5)

Two frames and no more, because at half-frame spacing a given output sample only ever sits under two windows. Now count the total weight arriving at , and this is where the two applications of show up together. Both halves already carry one factor of inside them, inherited from in (3.1), and (3.5) supplies the second one explicitly. So the tail arrives weighted by and the head by . For the output to be the signal itself, and not the signal with a tremolo on it at 250 Hz, those two have to come to exactly one:

(3.6)

That is a real constraint on which window you are allowed to pick, and the square root of a Hann window is probably chosen precisely because it satisfies it (probably a staple choice in DSP world, that me personally is not familiar with). Squaring the square root gives back a plain Hann window, and the two halves of a Hann window are cosines exactly half a period apart:

The cosines cancel, the two terms come to , and (3.6) holds at every for free. So the tapering undoes itself exactly, the joins vanish, and what emerges is the signal with the gain applied and no seams in it. That is (windowed )overlap-add, and any stage that rewrites a spectrum bin by bin ends up needing it.

Some of the pieces that feed that gain:

Residual echo + reverb model

The linear filter is a model of a linear path, and rooms are not perfectly linear - clipped speakers, vibrating laptop chassis, and a reverberant tail longer than the filter's own 52 ms all leak through. AEC3 estimates the residual echo's power per FFT bin, adds a reverb tail on top, and computes a per-bin suppression gain in . Frequencies that are mostly echo get reduced but frequencies that are mostly near-end speech are left alone. That gain is applied to the linear output before the inverse FFT, or to the raw capture spectrum instead whenever the state machine has decided the linear filter isn't currently trustworthy.

The reverb tail is a one-line recursion per bin:

(3.7)

It is tempting to call this a learned reverb model, and that turns out to be only half right, in the direction you would probably not guess. , the spectral shape of the tail is estimated, pulled out of the linear filter's own frequency response as it converges, whereas the decay just sits at the constant 0.83 on default config, with the whole linear-regression apparatus that exists to estimate it hiding behind a flag that only trips if you set the default decay negative.

Near-end detector + dominant-nearend mode

The single hardest moment for any AEC is when both sides are talking - the so-called double-talk condition. Pushing the suppression gain too hard in that moment chews up your voice. AEC3 ships two detectors for this and picks one at construction, and the one you get by default is the dominant-nearend detector, which is a good deal cruder than the name suggests. It sums energy over bins 1 to 15 only, which at a 16 kHz band rate with 65 bins works out to roughly 125 Hz through 1.9 kHz, and then asks two questions of that single low band: is the residual echo below a quarter of the near-end energy, and is the near-end itself well clear of the comfort-noise floor? Both have to hold for 12 consecutive blocks before it latches, and once latched it holds for 50 blocks unless echo comes back hard enough to knock it out early, which is the hysteresis that keeps the whole thing from flickering mid-syllable.

When it does latch, the suppressor swaps its whole parameter set for a more permissive one, and the threshold at which a bin counts as transparent enough to leave alone goes from 0.3 to 1.09, so roughly three and a half times more echo gets tolerated before the gain starts biting. That swap of a table of constants is the entire "pull back on suppression to protect the near-end voice" mechanism.

The subband detector compares two configurable frequency regions against each other, and it ships disabled. If you ever go to turn it on, look at the regions first, because both of them default to the single bin 1 and it would sit there comparing a band against itself.

Transparent mode

If you're wearing headphones, there is usually no acoustic echo path because the speaker cannot reach the mic. A linear canceller subtracting "almost nothing" is harmless, but a suppressor is different: it applies frequency-dependent gains, and those gains can dull, gate, or otherwise colour the near-end voice even when there was no echo to remove. So AEC3 tries to detect the no-echo case and get out of the way.

There are two classifiers here too, and they have about as little in common as the two update gains did. The default is a pile of heuristics watching how well the main filter is converging, how long it has gone without converging while render was active, how many blocks it has spent fully diverged, and whether it has ever locked onto a sane short delay, with the rough summary being that if the filter has had six seconds of strong un-saturated render and still has nothing to show for it then there is probably no echo path to find. The alternative is a two-state hidden Markov model, off by default, and it ignores every one of those signals in favour of a single bit per block, namely whether the shadow filter cleared a relaxed convergence threshold, which it then runs through transition probabilities with a dead zone between activating at 0.95 and deactivating at 0.5.

The way transparent mode takes effect is indirect and I think rather elegant. Instead of touching the suppression gain, it reaches into the residual echo estimator and drops the assumed echo-path gain to 0.01 in amplitude, a factor of ten thousand in power, while skipping the reverb contribution entirely. The suppressor then goes about its ordinary business on an echo estimate that has collapsed to nearly nothing and settles on a gain of about 1 all by itself, with nothing anywhere in the code having to special-case the headphone situation.

Comfort noise

A suppression gain that swings between 1.0 and 0.05 several times per second sounds horrible - the noise floor pumps in and out. So AEC3 backfills the gated bins with a small amount of shaped comfort noise, and the weighting falls out of one line:

(3.8)

is a noise spectrum shaped to the estimated capture noise floor with randomised phase, and the weighting means the noise arrives exactly where the gain took something away: a bin left alone at gets none, a bin gated to gets all of it. So it is a power-complementary crossfade between the suppressed signal and synthetic room tone, per bin and per block, and it happens in the same pass as the gain itself and before the inverse FFT. The silence in between speech then sounds like a room and not like a software product.

None of these is the "real" echo canceller on its own. The trick is that all of them are running at the same time, on the same 4 ms blocks, and mostly talking to each other through one shared block of state that tracks delay confidence, filter convergence, saturation, and the reverb and ERLE estimates.

The cadences are easy to mix up as well, so to be explicit: the API takes 10 ms frames of 160 samples at the band rate, while everything described in this post runs on 64-sample blocks of 4 ms. Those numbers do not divide, and 160 is not a whole number of blocks, which raises a fair question about where the leftovers go.

They get buffered, and none of them are thrown away, as in: each 10 ms frame is handed over as two 80-sample sub-frames, and each sub-frame extraction takes whatever is already buffered, tops it up to 64 from the incoming samples, emits that as a block, and keeps the rest. Run the arithmetic forward from empty and the buffer holds 16 samples, then 32, then 48, then 64, and when it reaches a full 64 an extra block is pulled out and the buffer empties. So the first 10 ms frame produces two blocks, the second produces three, and it alternates like that forever. Over 20 ms that is 5 blocks of 64, which is 320 samples for the 320 that came in, exactly.

That balance does not come for free, and the cost is a fixed delay. Of course, thats because handing back 160 samples on every call while only ever receiving whole blocks means the output side has to begin with something already in hand, and it does: the re-framer that turns 64-sample blocks back into 80-sample sub-frames starts life holding a block of 64 zeros. Those zeros are the first thing you get back, and every real sample afterwards comes out 64 positions late. That is 4 ms, fixed and permanent.

It is a nice detail that this is 64 and not the 32 the leftover count might lead you to guess. The extra block only appears at the end of every second frame, and by the time it does, the output side has handed out 320 samples while having received just 256. It needs a whole block in reserve to cover that gap, so a whole block is what it starts with.

The suppressor then adds a second 4 ms for an unrelated reason. (3.5) cannot finish an output block until the following frame has been transformed, because it needs that frame's rising head to add to this one's falling tail (as explained previously in the post), so its output permanently trails its input by one block. Stack the two and the framing and suppression path costs 8 ms, or 128 samples, end to end. What you get for that price is exact reconstruction: outside the startup transient, the samples coming out are the samples that went in, shifted by 128 and otherwise untouched apart from the gain you asked for. The three-band split and merge add delay of their own on top, which I have not measured.

One consequence of all that is worth knowing if you ever use the linear output: it never touches the suppressor, so nothing in it passes through (3.5), and it pays the framing delay without the overlap cost. That puts it at 4 ms where the processed output sits at 8. The two are handed back through separate re-framers and they are not aligned with each other, the linear one running a whole block ahead. Every decision in this post is made once per block.

At the end of the day, the output you hear on a call is the linear filter doing most of the work and small augmenting estimators tidying it up to sound "right".

Notes

  1. Block size at the band rate is the same as the half-FFT length, which is why everything fits together without zero-padding gymnastics. AEC3 runs at 250 blocks per second.
  2. The matched-filter window length and inter-filter shift come from MATCHED_FILTER_WINDOW_SIZE_SUB_BLOCKS = 32 and MATCHED_FILTER_ALIGNMENT_SHIFT_SIZE_SUB_BLOCKS = 24 in the reference implementation.
  3. "Decimation factor" in this post means the matched filter's downsampling factor, not the three-band split - those are two different downsampling steps stacked on top of each other. At 48 kHz with the default factor of 4 you are going 48 -> 16 kHz by band split, then 16 -> 4 kHz by decimation, and only the second number moves if you change the config.
  4. Every default quoted in this post (13 partitions, 5 matched filters, factor 4, 0.83 decay, 5 and 20 votes, and so on) is the shipped config. Nearly all of it is tunable, and a few of the behaviours described here flip entirely if you tune it - the reverb decay and both of the detectors in §03 being the obvious ones.
  5. Reference AEC3 library (in Rust): https://github.com/RubyBit/aec3-rs

That's all from me today. I think it's quite interesting that such an intricate piece of software is running on our every call and no one pays mind to the complicated inner workings of it! Kudos to the WebRTC team for this amazing open source version of a versatile acoustic echo cancel algorithm.

- END OF SHEET 01 -

Drawing Angelos · Sheet 01 Rev. A · 21 May 2026 · do not scale from print