)
Series 2 · Neural Networks

Why Neural Networks Are So Powerful — and So Opaque

A large enough neural network can in theory approximate any relationship, but the theorem that guarantees this says nothing about how to find it; and the weights training does find work without anyone knowing what they mean. It is from this opacity that mechanistic interpretability was born, with tools such as sparse autoencoders and circuits.

← All topics
The power of neural networks and the explainability problem: the universal approximation theorem on one side, the impossibility of accounting for why an output was produced on the other.

In chapter 6 we said that by combining enough neurons we can represent any relationship between input and output — and we added that this was a precise mathematical result, though one carrying an important caveat we put off until later. It is time to settle that debt.

The theorem, and what it does not say. The result is called the universal approximation theorem: a network with at least one hidden layer, of sufficient size and with a non-linear activation, can approximate any continuous function on a compact domain to whatever precision we like. Put that way it sounds like an unlimited promise, and this is where the caveat is needed: the theorem guarantees that such a network exists. It does not tell us how to find it, nor how large it would have to be, nor whether gradient descent will ever reach it in reasonable time. It is a theoretical existence, not a practical recipe.

And it is a distinction worth keeping in mind every time someone says a neural network “can do anything”. It can, in principle. That it will actually get there in practice is guaranteed by no theorem: it is guaranteed only by the data, the architecture, and the optimization — that is, by all the work of the previous chapters.

The other face: opacity. Chapter 8 closed by observing that backpropagation finds weights that work, but tells us nothing about what each one means. Now we can say why. Every weight is a number that only makes sense in relation to millions of others; there is no dictionary translating a weight into a human concept; and the result is that not even the people who build the model can explain why, given a certain input, a certain answer came out.

This lack of transparency is not a theoretical annoyance: it is a concrete reliability problem. If we don’t know why a model answers the way it does, it is hard to trust its answers blindly in sensitive contexts — and those are exactly the contexts in which these models are being put to work today. It is from that concern that an entire research field was born, dedicated to understanding what happens inside the network: mechanistic interpretability.

Activation patterns and features. To see how it works, one concept has to be made precise first. Take a single layer of the network, with some number of neurons. Given an input, we can look at which of those neurons return a value appreciably different from zero: those neurons are, so to speak, on, and the set of all the ones that are on for a given input is its activation pattern — a specific point in a space with as many dimensions as that layer has neurons. Note that the dimensions are the neurons: there is nothing more exotic about them.

And the combination of neurons that light up together to encode a concept is what interpretability calls a feature. A feature does not live inside a single neuron: it is a direction within the activation space.

Superposition. Take the concept of the Golden Gate Bridge. We would expect that, from reading enough text about that bridge, the network would dedicate a neuron of its own to it. It doesn’t: that same neuron also fires on concepts that have nothing to do with the bridge. The reason is simple and structural — the features inside a network vastly outnumber the neurons available, so each one settles for a combination of several neurons rather than having one to itself. Different features end up sharing the same few neurons, overlapping: this is the phenomenon known as superposition.

To picture it, imagine a layer with just three neurons: a tiny space, barely three dimensions. Inside a space that small, hundreds of different concepts have to coexist. There is no room to keep them apart, so their directions end up crowded together, almost glued to one another: change the combination of those three numbers slightly and we slide from one concept to the next. It is a dense, crowded space, and hard for a human to read.

On the left, the activation space of a three-neuron layer with concept directions piled on top of each other; on the right, the space expanded by the sparse autoencoder, where the Golden Gate Bridge feature has a direction of its own, with Alcatraz and Vertigo nearby but distinct.
From a crowded space to a readable one. The sparse autoencoder does not compress: it re-expands a layer's few overcrowded dimensions into a much wider space, where each feature gets a direction more of its own. Experiment on Claude 3 Sonnet, Templeton, A. et al. (2024), Scaling Monosemanticity, Anthropic.

Sparse autoencoders. These were born to untangle the knot. A traditional autoencoder compresses a rich input into a few essential dimensions and then tries to reconstruct it. A sparse autoencoder does the opposite: it takes those few overcrowded dimensions and re-expands them into a much wider space, where every feature can get a direction more of its own, further from the others. The space becomes less dense, more spread out, and in that spreading each direction becomes more amenable to interpretation.

Not automatically legible, though: understanding what a given direction actually stands for still takes patient work, examining which inputs make it fire. That is precisely the work Anthropic’s researchers did on Claude 3 Sonnet, when they managed to pin down the direction tied to the Golden Gate Bridge. By forcing that single feature to stay active — a technique called steering — they made the model pull almost every answer back to the bridge, to the point of describing itself as the bridge. Around that feature they also found the neighboring ones: Alcatraz, Ghirardelli Square, Hitchcock’s film Vertigo.

One figure gives the measure of how far superposition is the norm rather than the exception: for 82% of the features extracted, there is no strongly correlated neuron at all. Features live in the combinations, not in the individual nodes.

We will return to dense and sparse spaces in the coming chapters, when we come to embeddings — where sparse will carry a different, almost opposite meaning to the one it has here: there it will denote a very high-dimensional representation devoid of meaning, to be compressed into a dense space in order to gain it. The two senses are worth keeping apart.

Circuits. The next step, and where research is moving now, is that of circuits. If sparse autoencoders tell us which features exist inside the network, circuits try to reconstruct the sequence in which those features switch on, one after another, layer after layer, until the final output — rather like reconstructing the wiring diagram behind a single piece of the model’s reasoning.

Neural networks remain opaque. But it is an opacity on which a great deal of ground is being covered today. And before moving on to language, one question of inventory remains: the fully connected network we have talked about so far is only one of the possible shapes. What are the others, and what are they for? That is the subject of the next chapter.