Lesson 7
FM
Additive charged an oscillator per partial. This lesson gets several dozen partials from two oscillators, by exploiting something you have already seen happen by accident. It is the cheapest method in this course by a wide margin, and the hardest to steer.
You have already built this
In lesson 4 we connected an oscillator to a parameter and dragged its rate upward. Past about 20 Hz the wobble stopped being a wobble and partials appeared either side of the carrier. We noted the implication at the time and moved on. This is that implication.
If modulating a frequency fast enough puts new partials into a spectrum, then modulation is not only a way to animate a sound. It is a way to build one, and an extremely cheap way, because the number of partials produced has almost nothing to do with the number of oscillators paying for them.
John Chowning worked this out at Stanford around 1967 while doing something else entirely, published it in 1973, and Yamaha licensed it into the DX7, which went on to become one of the best-selling synthesizers ever made.
Two controls, and they do not overlap
Two-operator FM has one oscillator, the modulator, driving the frequency of another, the carrier. Everything you can do with it comes down to two numbers.
The ratio between the two frequencies decides where the partials land. Sidebands appear at the carrier plus and minus whole multiples of the modulator frequency, so if the two are in a simple whole-number ratio, every sideband falls on a multiple of some common fundamental and the result has a definite pitch. If the ratio is not a simple fraction, they scatter, and you get metal.
The index decides how many partials there are. It is the peak frequency deviation divided by the modulator frequency, and turning it up pushes energy outward from the carrier into more and more distant sidebands. Index is brightness.
That separation is unusually clean: one control for pitch character, one for brightness, and they barely interact. It is also the only part of FM that is intuitive.
The part with a closed-form answer
FM is the only method in this course whose spectrum can be written down exactly. The amplitude of the nth sideband is Jₙ(index), a Bessel function of the first kind, and that is the exact answer rather than an approximation or a rule of thumb.
Which means the figure below can do something none of the others can. The gold stems are computed from Bessel functions with no reference to the audio at all, drawn before anything is played. The violet trace underneath is the analyser measuring what the oscillators actually did. They should coincide.
Predicted, then measured
Gold is what Bessel says. Violet is what the oscillators do.
Sidebands land on every harmonic of the carrier. The most ordinary FM sound there is, and the closest to a sawtooth.
Predicted spectrum
Measured spectrum
Press play to see the trace
- Carrier
- 220 Hz
- Modulator
- 220 Hz
- Partials above 0.4%
- 6
- Spectrum
- harmonic
Two oscillators are producing 6 partials. Additive would have charged you an oscillator each. The index alone decides how many: Carson's rule puts almost all the power inside 2(index + 1) × modulator, which here is 1,320 Hz wide, and you can watch that width grow as you turn it up.
Two things in that figure worth stopping on
Partials disappear at particular settings. Select the 1 : 3 ratio, then turn the index slowly and watch the line at 220 Hz. At an index of about 2.405 it goes to nothing at all. That is the first zero of J₀, and it means the fundamental of an FM sound is not guaranteed to be present, which is not a fact any synthesizer panel would ever tell you.
The spectrum is often lopsided. With a low carrier and a high index, carrier − n×modulator goes below zero. Those components do not vanish: they reflect back to their absolute value with their phase inverted, and can cancel against sidebands already sitting there. The prediction above folds them deliberately, because predicting only the positive side would draw stems where the measurement shows nothing.
Those two facts collide, which is why the first one specified a ratio. At 1 : 1 the sideband two steps below the carrier sits at 220 − 2×220 = −220, folds back onto 220 exactly, and fills the hole that J₀ just made: the line there is not J₀ but J₀ − J₂, which at an index of 2.405 is 0.43 rather than zero. The same happens at 1 : 2. A fold lands on the carrier whenever 2 ÷ ratio is a whole number, so the null is only visible at ratios where it is not - which is 1 : 3, 2 : 3, 1 : √2 and 1 : 3.53 of the six on offer.
One implementation detail that matters enough to state. Real FM adds its deviation in hertz, and the Bessel result only holds if it does. Lesson 4 modulated pitch in cents instead, because a musical vibrato should be the same size at every pitch, and cents are logarithmic. That is the right choice there and the wrong one here, and it is exactly why lesson 4's spectrum smeared at high depth and had to be limited. Same two oscillators, different arithmetic, different maths entirely.
Where it becomes an instrument
A steady FM tone is a curiosity. What makes it an instrument is putting an envelope on the index, and this is the idea worth taking away from the lesson.
A high index at the attack means many sidebands, so the note starts bright. Let the index fall as the note decays and the sidebands leave, so the sound darkens by itself. That is precisely what a filter envelope does in subtractive synthesis, achieved with no filter, by never generating the partials in the first place rather than removing them afterwards.
It is also why FM was so good at the sounds it became famous for. Bells, electric pianos, tubular things, plucked metal: every one of them is bright at the attack and dark immediately after, which is a falling index almost by definition.
Four numbers, and it is an instrument
Strike each one. Then flatten the index envelope and strike again.
A harmonic ratio and an index that collapses. Bright bell-like attack, warm sustain. Four numbers, and it is the sound of most of the 1980s.
Press play to see the trace
- Ratio
- 1 : 1
- Index at attack
- 4.20
- Index after decay
- 0.40
- Oscillators
- 2
Every one of these is two oscillators. Drag the index envelope to flat and strike again: the note keeps its volume shape and loses its character entirely, because the brightness was never coming from the amplitude envelope. It was coming from partials appearing and then leaving.
So why is FM hard, and why did it stop
Everything above is two oscillators. The DX7 had six per voice, with 32 fixed ways of wiring them together called algorithms, each operator having its own envelope. That is enormously more powerful and it is where the trouble starts.
The parameters do not map onto perception. A cutoff knob does one thing and it does it monotonically: turn it up, get brighter. An FM ratio does not behave like that at all. Nudge it from 1 to 1.01 and a clean tone becomes a slow metallic beating. Nudge the index and partials do not merely get louder, they trade places, because Bessel functions oscillate. There is no direction to turn a knob in that reliably means more.
Which made it famously unprogrammable. The DX7 shipped with presets that most owners never got past, and a generation of records used the same few factory sounds because editing meant reasoning about six interacting Bessel spectra through a membrane keypad and a two-line display.
And then memory got cheap. FM won in 1983 because it produced complex evolving spectra for almost no computation, at a time when computation was the binding constraint. Once you could simply store a recording of a real bell, the argument for computing one collapsed, and by 1990 sample playback had taken the market.
It did not go away, though. FM is cheap, and cheap still matters where there are thousands of voices or a tiny power budget, so it survives in software instruments and in every game console sound chip of the era. What changed is that it stopped being the obvious answer and became one technique among several.