
Chapter thirteen closed by promising that the Transformer would solve the limits of Word2Vec: the same vector for the same word, regardless of its shades of meaning, and no sensitivity to word order within a sentence. But that wasn’t the first promise left hanging: back in chapter ten we had already inherited the fundamental limit of RNNs, sequential reading that causes distant dependencies to decay through the vanishing gradient. A single architecture solves both problems at once, and it’s time to open it up, piece by piece, for the rest of Series 4. Let’s start from where it was born.
In 2017, a group of researchers from Google, across Brain and Research, and from the University of Toronto published a paper with a title that’s already a manifesto: “Attention Is All You Need”. The problem they wanted to solve was machine translation, a specific case of what the literature calls sequence transduction: taking a sequence in one language and turning it into the corresponding sequence in another. Nobody, at that moment, was thinking about a chatbot. And yet that architecture turned out to be far more general than the original goal, to the point of becoming, within just a few years, the architecture virtually every large language model in use today is built on.
The heart of the idea is called attention, and it’s a paradigm shift from how an RNN works. A recurrent network reads one word at a time, carrying along a summary of what it has already read. The Transformer, instead, looks at the whole sequence at once and, for every word, directly computes how relevant each of the others is to it. If “French” depends on “Paris” ten words earlier, attention connects them in a single step, not through a chain of ten hidden states to climb back through.
There’s a useful way to picture how this architecture is built as a whole. The input is still a matrix, only instead of being the matrix of a flattened image, it’s a matrix where every row is a word, or more precisely a token, as we’ll see in the next chapter, and holds its vector representation, the embedding from chapters eleven and twelve: every column is one of the dimensions of that vector. That matrix then passes through a stack of identical blocks, dozens of them in modern models, and comes out of each one with the same dimensions it went in with. The hourglass shape from chapters ten and twelve doesn’t disappear entirely, though: we find it inside every block, in the attention layer. That layer is split into many heads working in parallel, and each one projects the vectors into a smaller space, from 512 to 64 dimensions in the original paper, before the results are recombined and brought back to the starting dimensions. The feedforward layer that follows does the opposite: first it widens, then it narrows. We’ll look at both up close in the next chapters.

And this is exactly where the literal meaning of the paper’s title comes from. Unlike a CNN, convolution disappears entirely here; and unlike an RNN, sequences of text are handled without any recurrent mechanism at all. Only attention is left, doing all the work that used to be handled by convolutional filters or by hidden states passed along over time: attention is all you need.
Looking at all positions together has a second consequence, less visible but just as decisive: it’s a computation that parallelizes on a GPU in a way an RNN, inherently sequential, simply cannot. It’s exactly this parallelization that made it possible to train models on amounts of text that, with a recurrent architecture, would have remained out of reach.
Back to the two limits of Word2Vec. The first was order: a sum of context vectors is commutative; it carries no positional information at all. The Transformer inherits that same problem at the root, and solves it by explicitly adding positional information to every word, the so-called positional encoding, which we’ll come back to later. The second limit was that the embedding algorithm always assigns the same word, meant as a sequence of letters, the same, single vector, regardless of the meaning it takes on depending on context. “Bank” stayed the same vector whether the sentence was about a financial institution or the edge of a river. With attention, every occurrence of a word dynamically weighs the words around it, so its representation shifts sentence by sentence.
One last clarification: the 2017 Transformer had two distinct halves, an encoder and a decoder. Modern large language models keep only the second half: they’re decoder-only models. A word of caution, though: in the context of Transformers, “decoder” doesn’t refer to the same-named block seen in the autoencoder from chapter ten, the half that widens the latent space back out to the original dimensions. In the Transformer the blocks neither narrow nor widen: the word “decoder” only indicates which of the two blocks of the original Transformer is kept, the one that generates text one word at a time, looking only at the words that come before it.
Over the next chapters we’ll take apart a decoder-only Transformer, layer by layer. We’ll do it by imagining feeding this architecture the sentence “The monkey didn’t climb the tree”, adapted from an example in the post with which Google introduced the Transformer in 2017, and we’ll discover how the model produces the output “because it was too lazy”.
