While music can be a hard thing to define, I like the simple definition of ‘organised sound’, because it suggests a structure, the presence of patterns and regularity, a means of forming expectations.
Whether it’s complex jazz or a simple children’s melody, there is always some structure involved. A rhythm, a phrase, or a key — an element that for the keen mind can be found, followed, and forecast.
Until it can’t. Until the direction changes and your prediction is broken, and the mind scrambles to grasp the new structure.
That’s what keeps it interesting, and I’d argue it’s the most significant element in our aesthetic enjoyment of music — the push and pull of familiarity and novelty, of uncertainty and resolution.
This fits in well with a prominent theory of cognition called ‘predictive coding’. Predictions and prediction errors are what the brain deals in, and it’s what makes music pleasurable.
The Brain as a Prediction Engine
The theory goes back at least at least as far as the 1860s, with the physicist Hermann von Helmholtz and his idea of unconscious inference.
“He was […] interested in theories of perception and argued that we perceive the world only thanks to a kind of unconscious reasoning or inference in which the brain is asking itself, “Given everything I know, how must the world be for me to be receiving the pattern of signals currently present?” This is the question that perceptual systems are built to resolve.” —Andy Clark, The Experience Machine

The modern version emphasises that the brain is constantly predicting its sensory inputs based on mental models of the world around it. When the prediction doesn’t match the signal, an error occurs, an adjustment is necessary.
All of perception, cognition, emotion, and action are thought to work together to minimise these prediction errors. Everything from vision and hunger to pain and even consciousness are suspected to be built on these predictions and the error correction process.
This is in contrast to the view that our conscious experience is a passive, mirror-image of our environment, something that must come in from the outside before we determine what’s there. This order is somewhat backwards.
We predict, based on learned models, and those predictions are checked against sensory evidence. This cascade of top-down predictions shapes our experience.

It’s a subtle difference, and the constant dynamic interplay between predictions and the senses makes the distinction unnoticeable in most of our experience. But as we’ll see, music and other examples support the predictive theory.
“[…] we are constantly predicting the future and hypothesizing what we will experience. This expectation influences what we actually perceive. Predicting the future is actually the primary reason that we have a brain.” — Ray Kurzweil, How to Create a Mind
Experiencing Predictions
The theory helps explain a lot of odd experiences that don’t fit in the traditional view that the mind is a reflection of the world.
Visual illusions are the most striking example. Based on certain shapes and colours, your brain expects to see movement — and you see that expectation.
Here are a couple of examples from Akiyoshi Kitaoka:
This is not seeing reality as it is.
A similar phenomenon occurs in auditory illusions, like the McGurk effect, where what we hear is influenced by the visual cues we see; or the phonemic restoration effect, where our brain fills in missing sounds in speech.
Consider an ambiguous recording that made its way across the internet. It’s a degraded sound but it’s possible to hear something, and what you hear depends on what you expect. If you read ‘green needle’ as you listen, you hear that; if you read ‘brainstorm’ as you listen, you’ll hear that instead. You don’t even need to read, just think the words and you’ll likely hear them.
The theory also helps explain the well-documented placebo and nocebo effects, where our expectations of a treatment produce the result, despite the treatment being inert.
The predictive coding theory aims to explain all perception, cognition, and action. For this to work well, the predictions need to be based on accurate mental models, and effectively adapt to the errors from mismatches with sensory evidence.
As these examples illustrate, we tend to experience the brain’s predictions, but illusions aside, most of the time they’re reliable. Having picked up the important patterns and regularities in our environments, our internal models are accurate and useful enough to make good guesses about what’s happening outside the dark confines of the skull.
When it comes to music, our models and predictions are put to the test, taken on a ride of expectation and deviation that renders the whole experience quite enjoyable.
Hierarchy of Patterns
Music is organised sound, and it’s organised over many different levels.
Take a single note — it’s a stable frequency over time, such as 440 Hz (the A note). A simple, predictable pattern. If that’s all you hear, you’d be listening to a ‘pure tone,’ not very exciting.

Nature tends not to produce pure tones — you’ll need a tuning fork, synthesiser, or oscillator for that. When you hear wind in the trees, running water, bird calls, or someone whistling, you get a wide range of overlapping frequencies.
Most musical instruments also produce complex combinations of frequencies, known as overtones, that follow a harmonic pattern. They’re integer multiples of the ‘fundamental frequency’ — ie: if 400 Hz is the fundamental, you’ll also hear 800 Hz, 1,200 Hz, 1,600 Hz, and so on.
Depending on the instrument, the overtones will play with different strengths and change over time, which is what gives the instruments their characteristic sound, it’s why the A note on a guitar doesn’t sound like the A note on the piano.

Given the overtone series is unique to each fundamental frequency, you can always determine the fundamental even when you only have the overtones. You could figure it out by doing the maths, but your brain seems to understand this on an intuitive level.
“The brain is so attuned to the overtone series that if we encounter a sound that has all of the components except the fundamental, the brain fills it in for us in a phenomenon called restoration of the missing fundamental.” —Daniel Levitin, This Is Your Brain on Music
The effect’s been used by producers to add an extra sense of bass. You might have speakers that can’t produce the lower notes, yet by a trick of human psychology still get a sense of the lower end.
The overtone series is just the beginning, it’s what you get when you strike one note on your guitar. Start playing chords, melodies, and rhythms, and you have even more patterns to pick up.
Stop for a minute and consider what hits your ears when you listen to a song. A very complex waveform, that looks something like this:

Your brain’s ability to pick out individual instruments from an orchestra or band is an incredible feat of pattern recognition (or prediction? We’ll return to that shortly).
You have the luxury of simply hearing the guitar, just by focusing on it. Your brain pulls out all the appropriate frequencies from that mess of waves that correspond to the desired instrument.
Throw in an understanding of how genres and styles sound, how cultures and different traditions compose, and how particular musicians sound and play — now you have an intricate hierarchy of patterns stacked upon patterns upon patterns.
All ripe for the neural machinery to not just follow along to, but to try and stay one step ahead of.
Follow Along or Stay Ahead?
Hold on, how do we know we predict these patterns and not simply follow them? Filling in the missing fundamental frequency requires predicting the nonexistent bass note, but that’s not really staying a step ahead, that’s filling in missing pieces.
The easiest example would be to consider the way we dance or tap our feet to the rhythm — you can only do that well if you are predicting what comes next. You’re too late if you wait until you hear the beat to tap your foot.
The same can be said of humming along. Even when you listen to a new song, it only takes a moment before you get a hold of where the melody is going. You can’t hum with a delay, so you must be looking ahead.
Another example comes in the way composers toy with tension and resolution. In Western music theory, a ‘cadence’ marks the end of a musical phrase or piece, and provides a sense of resolution — it’s like coming home.
There are several different cadences, but one of the most recognisable is the perfect cadence, which moves from the dominant (V) to the tonic (I) chord, and creates a strong feeling of closure.
As music moves between chords, there’s a natural pull toward the tonic. Skilled musicians play with this expectation, leading us towards the resolution only to dart off in another direction, prolonging the tension.
Here’s a ‘deceptive cadence’:
These are both simple, unadorned examples, but try listening out for cadences in the music you listen to, you’ll find them often.
Beyond the types of expectations we’re familiar with from our own listening experiences, there’s also a lot of support for the predictive hypothesis in the science literature.
Music, they say, “provides a remarkable ‘epistemic offering’: it continuously satisfies our need to resolve uncertainty, thereby underwriting its rewarding nature,” and “music perception, action, emotion and learning all rest on the human brain’s fundamental capacity for prediction.”
A recent study also confirmed a hierarchical processing of sound through different areas of the brain, with feedback and feedforward connections likely representing the dynamic interplay between predictions and error signals.

So there is decent support for the brain as a predictive engine, that musical perception and enjoyment are dependent on these predictions, and that these predictions of music form a hierarchy, much like the music itself.
There is just one problem to sort out.
Minimising Musical Errors
One critique of the predictive processing theory is the ‘dark room problem’ — If the goal of the predictive brain is to minimise prediction error, wouldn’t the best thing be to sit in a dark room where no surprises can be found?
In terms of music, shouldn’t we listen to the most familiar, simplest song imaginable? Or just sit in silence? That would guarantee no errors.

Clearly we enjoy some novelty, complexity, and variation in our song choices. We have an intrinsic drive towards new experiences and learning new things. To do the same thing over and over is boring.
The predictive processing model does not mean that we — and organisms in general — are designed to avoid stimulation and change. We’re thrown into a dynamic world, faced with hidden risks and rewards, and we need to find a way to live in this environment.
A creature that can minimise surprise not in a dark room, but in a complex and risky environment, is more likely to pass on its genes. Making good guesses in such an environment requires good mental models, and those need to be shaped through experience.
So curious creatures — if they don’t die in the process — will end up with better models and a predictive advantage.
When we listen to music, minimising prediction errors is part of the fun — avoiding them altogether is not. There’s no risk of death and destruction if you get them wrong, you don’t need to be scared of the surprises, only of music you don’t enjoy. It’s a safe bet.
Music offers a complex, dynamic, largely risk-free way to learn new auditory models, update the old ones, and broaden our understanding, so as to make better guesses in the future.
Unfamiliarity in the Familiar
I want to take a moment to address why we can keep going back to the same music. There are some songs I imagine you have listened to hundreds of times, that you can keep going back to, and keep getting enjoyment from.
Why doesn’t that music get boring? How can there be any prediction errors left to minimise? Surely you know it inside and out!
In some cases, the music is tied to nostalgic memories and emotions, sometimes it’s just particularly good to dance to. The predictive processing theory might say we find comfort in the familiarity, it’s safe because we’ve been through it before and nothing bad happened—almost like a dark room, but we just finished criticising that.
These provide some of the story, but I don’t think they’re the whole picture. Allow me to speculate for a moment on how else the predictive processing theory might be involved. A couple of possibilities come to mind:
The Devil’s In The Details
First, despite your many listens, I would wager that there are still unknowns. This might be because you forget certain features of the song, or because you don’t really internalise all of it to begin with (I still don’t know the lyrics to songs I’ve heard numerous times).
When we say we’ve listened to a song 100 times, I’d wager not many people have listened to a song 100 times on repeat, over and over again. Even our most cherished songs are likely to gnaw at us then. Love them as we do, we need some time apart — to forget a little perhaps?
And what might you forget? Remember how complex music is — all those waveforms stacked upon each other, hierarchies of simple shapes layering together and changing over time. Look at how detailed music is as we zoom in from 6-minutes down to 50 milliseconds:

If you listen only to that 50 millisecond clip, you won’t hear much more than a faint click sound, but there’s a lot of information contained within. When you pass over that moment as you listen to the whole song, all those sharp peaks and waves support your perception of guitars and synths and drums and vocals.
How much of that detail do you store in your memory? I think we tend to remember more of the higher level stuff, melodies and chord progressions, lyrics and rhythms, the type of thing that lets you recognise the song even if it were played by a new band or in a different style.
But how accurately do you remember the particular tones of the instruments? The notes tapped during a guitar solo? The slight variations in the strength of the keys played on the piano? The timing of the drummers’ hi-hat during the chorus?
Music is rich, deep, full of sound that you wouldn’t notice unless it was changed or taken out. You might not consciously attend to these features, but they add to the experience. The predictive brain might still be challenged in the details of a familiar track, even more so if it hasn’t heard it in a while.
Crafting Expectations
The second reason I think we go back to familiar songs is that they are like stories — we can listen to them again and again because, despite their familiarity, they are expertly crafted to lure us back into the same old expectations.
A lot of songs rely on familiar chord progressions—take the 12-bar blues for a prominent example. This is a standard progression that’s provided the backbone to countless songs across genres. Progressions are reused repeatedly in part because they’re so effective at taking us on a journey.
Or consider that when we’ve learned certain musical patterns, we need to rely on convenient starting points to jog our memory of the latter parts.
For example, it’s easy to run through the alphabet with the melody you learned as a child, but jumping straight in with ‘K’ is difficult. It helps to find a natural starting point of a phrase, like H-I-J-K or L-M-N-O-P.

Later parts of a familiar melody still need to be triggered by the preceding steps. The elements aren’t able to be recalled like a computer pulling up data, they’re sequential patterns that need to be played out.
This all suggests to me that when we hear the beginning of a melody, our prediction engine gets to work, forming expectations about the next note or phrase.
Those expectations might play out, or they can be dashed — and for familiar tracks, we know this in advance — but we can’t help but get caught up in the process. It’s like watching a movie or reading a book for the 2nd or 3rd time. We know who dies, we know the ending, but we’re on the ride again and holding out hope.
You can’t stop the visual illusions from tricking your brain, and we can’t help but get excited for the bass drop or the guitar solo, even though we’ve heard it all before. It all tempts the mind back into that giddy sense of anticipation.
The Ear of the Beholder
“Nothing we do or experience — if the theory is on track — is untouched by our own expectations. Instead, there is a constant give-and-take in which what we experience reflects not just what the world is currently telling us, but what we — consciously or nonconsciously — were expecting it to be telling us.” — Andy Clarke, The Experience Machine
We don’t see reality as it is, but as we expect it to be. Our predictions shape our experience, our experience shapes our mental models, which shape our predictions. The process reinforces and feeds back on itself.
It’s not a passive form of reception, but an active system that generates and tests theories. Sometimes these tests fail, the predictions come up short, and the errors cascade up the hierarchy, adjusting the models along the way, leading to new theories and predictions.
Music is organised sound, organised in such a way as to tease this predictive machinery. You must step in time, you must move when it moves, you must anticipate, and adapt when it dashes your expectations.
It operates over many levels, from the simplest fluctuations of an A note, through the unique sound of instruments and voices, to the melodies, rhythms, progressions, lyrics, and styles.
It takes us on a journey, luring us down a path of highs and lows, of tension and resolution, expectation and deviation. In a matter of seconds, music can get you bobbing your head, tapping a foot, singing along, dancing with others, playing air guitar, or just scratching your head.
“… the properties of musical movements which possess a graceful, dallying, or a heavy, forced, a dull, or a powerful, a quiet, or excited character, and so on, evidently chiefly depend on psychological action.” —Hermann von Helmholtz, On the Sensations of Tone

Be First to Comment