
Chapter 10 closed on an uncomfortable constraint: a neural network, whatever its topology, speaks only the language of numbers. If we want a model to read a sentence, that sentence must first become numbers. Assigning a number to a word is, in itself, a trivial task: a simple list will do — we will use one in a moment, precisely to see its limit. The real problem is different: we want that number not to be just any label, but to carry the word’s meaning with it.
But what does it actually mean for a number to carry meaning? Here it helps to take a step back and ask what meaning itself is. When we try to define a word, we almost always do it with another word, a synonym: meaning is never something absolute, standing alone, but a relative concept. Two words are similar in meaning when they are, in some sense, close. And if meaning is a relationship of closeness, then we can translate it into something we know how to handle mathematically: geometric closeness. The guiding idea behind all numerical representation of language will be exactly this: build a space in which words with similar meaning occupy nearby positions.
Let us try the most obvious approach. Take a simplified vocabulary where the first word is Abacus and the last is Zoom, and assign each word its own position in alphabetical order: a single number, on a single line. Does it work? Not at all. In this representation Dog ends up closer to Abacus than to Mammal, and Mammal closer to Mussel than to Dog. Position in a sequence reflects the order in which we chose to list the words, not the closeness of their meaning.
There is a second limit, an even deeper one, that no reordering of the list can fix. Even choosing the most sensible possible criterion for lining the words up on that single line, it would still be impossible to represent the fact that a word is equidistant, in meaning, from several different words at once: Dog, for instance, should be able to sit as close to Mammal as to Friend and to Pet. But laid out on a line, these words will inevitably end up at distances from one another that do not reflect the distances between their meanings. What is needed is more space, literally: we have to move from a linear representation, on a single axis, to a multidimensional one. In a space with more dimensions, every word becomes a vector, and each component of that vector is a coordinate on one of the space’s axes. The simplest vector representation of this kind is one-hot encoding.

Every word in the vocabulary becomes a vector as long as the entire vocabulary, made up entirely of zeros except for a single one in the position that identifies it. But the resulting space is sparse, because every component of each vector is zero except one. This representation encodes every word in a space with as many dimensions as there are words in the vocabulary, where each word lies on a different dimension, orthogonal to all the others: no two words ever have their one in the same position, and the result is that every word ends up at exactly the same distance from every other, whether it is a close synonym or a term with no relation in meaning at all.
The solution is to move from a sparse space to a dense one: instead of a vector as long as the entire vocabulary, each word is represented with a much smaller number of dimensions, each of which captures a shade of meaning rather than simply identifying the word. To get a sense of it, imagine a space of just three dimensions, which we will call Wings, Engine, and Sky. Helicopter would have a component on Sky and one on Engine, because it flies, but thanks to an engine, not wings. Eagle would have a component on Sky and one on Wings, because it flies thanks to wings, with no engine. Hydrofoil would have a component on Wings and one on Engine, because it has both but stays on the water, never touching the sky. Aircraft, which has all three features together, would instead have a component on each of the three axes. In this example we chose the meaning of the three dimensions ourselves, to make the mechanism plain, but in reality nobody decides that in advance: it is something the model learns on its own during training, by observing the contexts each word appears in: words that recur in similar contexts end up in nearby positions. And the dimensions that come out of it, as a rule, carry no readable label like Wings or Sky: they work, but taken one at a time they do not correspond to a concept we would know how to name. A dense representation of this kind, learned by a neural network, is called a word embedding, and it is exactly the subject of the next chapter, on Word2Vec.

