Lesson 1

Timbre

A flute and a violin playing the same written note are at the same frequency, the same loudness, and are instantly distinguishable by anyone. Whatever is doing the distinguishing is what this entire course is about controlling.

The thing left over

Take everything obvious away from a sound. Pitch is frequency, and we can set both instruments to 220 Hz. Loudness is amplitude, and we can match those too. Duration is duration. What is left, after every quantity with an obvious name has been fixed, is called timbre, and defining it as the leftovers is unsatisfying enough that it is worth doing properly.

It has exactly two ingredients, and this lesson exists because almost everybody guesses only the first one.

Ingredient one: which partials, and how loud

A pitched sound is never a single frequency. It is a fundamental plus a stack of partials above it, and in almost all pitched instruments those partials sit at integer multiples of the fundamental: 220, 440, 660, 880 and on up.

The instruments differ in how loud each of those is. That list of amplitudes is the sound's spectrum, and it is the first ingredient. A clarinet is nearly missing its even-numbered partials. A flute has very little above its fundamental at all. A bowed string has everything, falling off gently as you go up.

If you have read Foundations, this is lesson 2 of that course arriving from the other direction: there we wanted to know why the partials are there, here we want to know what to do with them. If you have not, nothing below depends on it.

Ingredient two: what happens over time

The second ingredient is the one people leave out, and it is arguably the stronger of the two. A sound has a shape in time: how fast it arrives, whether it holds, how it leaves. That shape is called an envelope, and it does more work in instrument recognition than the spectrum does.

The classic demonstration is a recording of a piano note with the first fraction of a second removed. It stops sounding like a piano almost entirely. The spectrum is untouched; only the arrival has gone, and the arrival was carrying most of the identity.

The figure below lets you run that experiment in both directions. Choose a spectrum, choose an envelope, and then change one while keeping the other fixed.

A spectrum and an envelope

Pick one of each. The interesting part is holding one fixed and changing the other.

Spectrum

Every harmonic at 1/n. The richest of the four, and the one lesson 2 shows you how to build.

Envelope

Energy arrives all at once and then only leaves. Nothing sustains it, so it decays from the first instant.

Press play to see the trace

Fundamental
220 Hz
Partials present
24
Spectral centroid
1398 Hz
Attack
4 ms

The spectral centroid is the amplitude-weighted mean of the partial frequencies, and it is the closest single number to what people mean by bright. The sine sits at 220 Hz because it has nowhere else to be; the string timbre sits at 1398 Hz, several times its own fundamental, which is why it sounds so much sharper despite being the same note.

What to actually listen for

Set the spectrum to String and play it Plucked. It sounds like something being plucked: a harpsichord, a pizzicato violin, a guitar. Now change nothing except the envelope, to Bowed. The same partials, in the same proportions, at the same pitch, and it is a completely different instrument.

Then hold the envelope still and change the spectrum instead. That difference is real too, but notice that it feels more like a change of tone than a change of instrument. Both ingredients matter; they do not matter in the same way.

The number under the display is the spectral centroid, the amplitude-weighted mean of the partial frequencies. It is the best single-number stand-in for what people mean by brightness, and it is worth watching because it makes an intuition quantitative: the sine sits at its own fundamental, the string timbre sits several times higher, and that ratio is roughly what your ear is reporting as brightness.

Two honest caveats

The envelope here is one shape applied to everything. Real instruments are not like that. On a real piano the high partials die away much faster than the low ones, so the spectrum is changing throughout the note rather than staying put behind a single fading curve. That is a per-partial envelope, and it is what lesson 6 is able to build once we can address partials individually.

Not every instrument has integer partials. Bells, gongs, drums and most struck metal have partials at ratios that are nothing like whole numbers, which is why they have a vague sense of pitch or none at all. That case is genuinely different and gets its own treatment in lesson 9, where we stop describing spectra and start simulating objects.

So a sound is a set of partial amplitudes and a shape in time. Every synthesizer ever built is a machine for specifying those two things, and the differences between the methods in the second half of this course are entirely differences in how they go about it.

Next we build the first ingredient from nothing. Lesson 2 is about oscillators, which are the standard answers to "which partials, and how loud", and about the arithmetic that makes those answers go badly wrong at high pitches.

Battuto is a free set of courses from Aphelion. We also make Phonon, a DAW built on everything in these lessons.