Almost every machine-learning task, at its core, reduces to the same mathematical problem: given a set of input variables, predict the value of the output variables.
Machine Learning FundamentalsIn traditional programming a human writes the rules; in machine learning you feed the machine examples and it works out the rules for itself.
Machine Learning FundamentalsA machine-learning system learns by adjusting its weights to make the loss — the number that measures the total prediction error — as small as possible, using an optimization algorithm, typically gradient descent.
Machine Learning FundamentalsA model that fits its training data perfectly can still be useless: the goal isn't zero error but generalization — predicting unseen cases as well as possible — which means resisting the temptation to overfit.
Machine Learning FundamentalsMachine learning can be traced back to three broad families — supervised, unsupervised, and reinforcement learning — set apart by what the training data looks like: labeled, unlabeled, or replaced by rewards.
A neural network is a model made of elementary nodes arranged in layers and joined by weighted connections: data enters through the input layer, passes through one or more hidden layers, and exits through the output layer — and the connection weights are exactly the parameters that training has to find.
Neural NetworksAn artificial neuron computes a weighted sum of its inputs, adds a bias, and passes the result through an activation function: it is the non-linearity of that function that makes depth meaningful, because without it an entire network would collapse into a single linear operation.
Neural NetworksBackpropagation computes the gradient of the loss by applying the chain rule layer by layer, from the output back toward the input: every weight receives its own share of responsibility for the error, and this is what makes it possible to train networks with billions of parameters, where solving the equation for the minimum would be unthinkable.
Neural NetworksA large enough neural network can in theory approximate any relationship, but the theorem that guarantees this says nothing about how to find it; and the weights training does find work without anyone knowing what they mean. It is from this opacity that mechanistic interpretability was born, with tools such as sparse autoencoders and circuits.
Neural NetworksThe feedforward, fully connected network seen in the previous chapters is only one of the possible topologies: CNNs, RNNs, Transformers, autoencoders, and GANs each exist to answer a different constraint imposed by the data — spatial, sequential, or the very absence of a label to predict.
A number only starts to carry meaning if it translates closeness of meaning between words into geometric closeness: the key to solving this problem is to use dense representations.
Language and MeaningWord2Vec learns dense word representations in a fully automatic way, training a neural network to predict a word from its context: the embedding we're after is a byproduct of that training.
Language and MeaningWord2Vec embeddings have structural limits, a single vector per word and no sensitivity to order, but they also hide surprising properties such as vector arithmetic and biases learned from text.
The Transformer, introduced in 2017 in "Attention Is All You Need", solves both the fundamental limit of RNNs and the two limits of Word2Vec at once: it looks at the whole sequence in parallel instead of reading it one step at a time, and it dynamically weighs context instead of assigning a fixed vector to each word.
Inside the Transformer