ISEGORIA / MATH ENCYCLOPEDIA
Room acoustics: the room is part of the speaker
What reaches your ears from a loudspeaker is only partly the loudspeaker. In the bass the room is a resonator whose standing waves decide which notes boom and which vanish, depending on where you sit; in the midrange and treble it is a cloud of reflections that decays at a rate set by the room’s volume and absorption; and a single nearby surface can cut notches into the response that no equaliser can fill. Three models, each exact within its assumptions, cover the three regimes.
Before you begin: the wave equation, standing waves and Fourier analysis
Predict, manipulate, then check your reasoning against the example and question. Graphs illustrate the mathematics; they do not replace a proof.
1. Room modes: standing waves in a box
Below a few hundred hertz the wavelength is comparable to the room, and the room stops behaving like open space: sound bounces between parallel walls and builds up standing waves at particular frequencies, the room’s modes. In a rectangular room with rigid walls each mode is labelled by three whole numbers, the number of half-wavelengths that fit along the length, the width and the height. The floor plan on the left is coloured by the pressure of the selected mode, oscillating: the two colours are opposite phases, and the pale bands between them are nodes, where that mode is silent. The loudspeaker sits in a corner, where every mode has a pressure maximum. Drag the listener: the response on the right is the sum of all the modes heard at that seat, and it swings by 10 to 20 dB from one frequency to the next. Above the Schroeder frequency the modes overlap so densely that the peaks and dips blur into a statistical average.
Worked example. In a room 5 m long, 4 m wide and 2.5 m high, the first axial modes are \(343/10=34.3\ \mathrm{Hz}\) along the length, \(42.9\ \mathrm{Hz}\) across the width and \(68.6\ \mathrm{Hz}\) from floor to ceiling. The second length mode is also \(68.6\ \mathrm{Hz}\), because the length is exactly twice the height, and the two modes pile up at one frequency. The first tangential mode, \((1,1,0)\), is \(\sqrt{34.3^2+42.9^2}=54.9\ \mathrm{Hz}\). A listener at the exact centre of the floor sits on a node of every mode with an odd \(n_x\) or \(n_y\), so the 34.3 Hz and 42.9 Hz modes are missing there. With \(V=50\ \mathrm{m^3}\) and \(T_{60}=0.4\ \mathrm s\), \(f_S\approx2000\sqrt{0.4/50}=179\ \mathrm{Hz}\).
Watch out. The walls are rigid and every mode decays at the same rate, set by a reverberation time of 0.4 s. Real walls flex and absorb more in the bass, which shifts and damps the modes differently. The source is a point in a floor corner and the listener’s ear is 1.2 m above the floor; the floor plan shows the selected mode at that height. Levels are relative, with 0 dB the average over the plotted range.
Why is a corner the loudest place for bass?
Every mode shape is a product of cosines, and each cosine is \(\pm1\) at a wall. At a corner, where three walls meet, every mode has its full amplitude: a source there drives all the modes and a listener there hears all of them. That is why a subwoofer in a corner plays loudest, and why pressure-driven bass absorbers, such as membrane traps, are placed in corners. The quietest bass seats are the opposite: points where many low modes have nodes, such as the middle of the room.
2. Reverberation: how long a room remembers
After a sound stops, its energy goes on bouncing from surface to surface, losing a fraction α at every reflection. Wallace Sabine, timing the decay of organ pipes in Harvard lecture rooms with a stopwatch and his ears, and changing the absorption by carrying in seat cushions, found that the time for the level to fall by 60 dB is proportional to the volume and inversely proportional to the total absorption \(A=\sum S\alpha\), measured in square metres of perfectly absorbing surface. Choose a room and its surfaces. On the left, the echogram of a handclap: the direct sound, a few distinct early reflections, then a dense tail whose level falls in a straight line on the decibel scale. On the right, the level of the direct sound and of the reverberant field against distance: closer than the critical distance you mostly hear the loudspeaker, further away mostly the room.
Worked example. A living room 5 × 4 × 2.5 m has \(V=50\ \mathrm{m^3}\) and \(S=2(20+12.5+10)=85\ \mathrm{m^2}\). A bare wooden floor (\(\alpha=0.10\)), plaster walls and ceiling (\(0.05\)) give \(A=2.0+2.25+1.0=5.25\ \mathrm{m^2}\); a sofa, curtains and two people add about \(6\ \mathrm{m^2}\), so \(A=11.25\ \mathrm{m^2}\) and \(T_{60}=0.161\times50/11.25=0.72\ \mathrm s\). A carpet (\(\alpha=0.30\)) adds \(4\ \mathrm{m^2}\) and brings it to \(0.53\ \mathrm s\). With a loudspeaker radiating into half-space, \(Q=2\), the critical distance before the carpet is \(\sqrt{2\times11.25/16\pi}=0.67\ \mathrm m\): at a normal listening distance, most of what you hear is the room.
Watch out. Sabine’s formula assumes a diffuse field, with the energy spread evenly and travelling in all directions. It fails in very absorbent rooms (it predicts a finite reverberation time even when α = 1, which Eyring’s formula corrects), in long or flat rooms, and in the bass, where the modes of the first experiment take over. Absorption coefficients depend strongly on frequency; the values here are rough mid-frequency figures. Air absorption, which shortens the treble decay in large halls, is left out. The early reflections in the echogram are drawn at random times with the right statistical density, not computed for this room.
Where does the 0.161 come from?
In a diffuse field the average distance sound travels between reflections is the mean free path \(4V/S\), so there are \(cS/4V\) reflections per second, each keeping a fraction \(1-\alpha\) of the energy. The energy decays as \(E(t)=E_0(1-\alpha)^{cSt/4V}\approx E_0e^{-cAt/4V}\), an exponential, which is a straight line in decibels. It has fallen by \(10^6\), 60 dB, when \(cAt/4V=\ln10^6\), so \(T_{60}=4\ln(10^6)\,V/(cA)=(55.26/343)\,V/A=0.161\,V/A\). Keeping \(\ln(1-\alpha)\) instead of approximating it by \(-\alpha\) gives Eyring’s formula.
3. The nearest surface: comb filtering and boundary gain
A hard surface near a loudspeaker acts as a mirror: the reflection reaches the listener as if from an image source behind it, a little later and a little quieter than the direct sound. Where the extra path Δ is a whole number of wavelengths the two add; where it is an odd number of half-wavelengths they cancel, so the response grows a row of evenly spaced notches, which on a linear frequency axis look like the teeth of a comb. On the left, the sound field at the chosen frequency, with the loudspeaker, its image and the two paths; drag the loudspeaker and the listener. Switch the surface to the wall behind the speaker and move the speaker close to it: the first notch climbs out of the bass, and below it the reflection arrives nearly in phase, adding up to 6 dB of bass. That lift is what the “near wall” setting on many active speakers is there to remove.
Worked example. A tweeter 1.0 m above a hard floor and an ear 1.2 m above it, 3 m apart: \(r_1=\sqrt{9+0.04}=3.007\ \mathrm m\), \(r_2=\sqrt{9+4.84}=3.720\ \mathrm m\), \(\Delta=0.714\ \mathrm m\). The first notch is at \(343/(2\times0.714)=240\ \mathrm{Hz}\) and the next at 721 Hz. With \(g=0.9\) the reflection arrives with \(0.9\times3.007/3.720=0.73\) of the direct amplitude and the notch is \(20\log_{10}(1-0.73)=-11\ \mathrm{dB}\) deep. For a speaker 0.5 m in front of the wall behind it, heard 2.5 m in front of it, \(\Delta=2\times0.5=1\ \mathrm m\) and the first notch is at 172 Hz; at 0.1 m from the wall it moves up to 858 Hz, and in the deep bass the response is lifted by \(20\log_{10}(1+0.9\times2.5/2.7)=5.3\ \mathrm{dB}\).
Watch out. One flat reflector and an omnidirectional point source. A real room has six surfaces, each adding its own comb, and a real speaker beams its treble forwards, so the reflection from the wall behind it is weaker at high frequencies than the model assumes; in the bass, where the model matters most, the source really is omnidirectional. The reflection coefficient is the same at every frequency here, whereas a carpet absorbs treble far more than bass.
Why can’t an equaliser fill in a notch?
The notch is a cancellation at the listener’s position between two copies of the same signal. Boosting that frequency boosts both copies equally, and they still cancel: the amplifier works harder, the rest of the room gets louder, and the notch stays. Only changing Δ, by moving the speaker, the listener or the surface, or weakening the reflection with an absorber, moves or fills it. Peaks are different: where the copies add, cutting the level works, and the broad bass lift near a wall is exactly the kind of peak a shelving filter can correct.
In practice
Where this mathematics and physics is at work, in explainers that take the real thing apart.