Episode 1's bowl, seen from the side. At the bottom, the best line: slope , still wrong by .
Still wrong.Less wrong.Still asking.
Local Minimum Pictures presents
a film by Learning Rate
A machine starts with two numbers and one question. Twenty-five episodes, two and a half centuries, one answer that keeps changing shape.
A drama series in twenty-five episodes, one every Sunday at 09:00 IST, from 8 Nov. Episode 1 airs Sun 8 Nov, 09:00 IST. Every number on screen is computed in your browser. Best on a desktop with a graphics card.
Part one · episodes 1–8
The learner gets a score and a way down. Then its first crisis: it can fool itself, so it learns to keep an honest exam.
Act one
A model is a few numbers. Learning is finding the numbers that make it least wrong.
Every Line Has a ScoreDownhillThe Line That LeansAsk the Neighbours
Episode 1 · Paris, 1805
“By the square of every miss.”
Linear regression and the loss surface
Sixty points, one line, two knobs. Every line you could draw has a score, and all of those scores together form a bowl whose bottom is the best line.
New ideas modelparameterspredictionresiduallossmean squared errorloss surfaceleast squaresconvex bowl
Episode 2 · Paris, 1847
“Less wrong, one step downhill.”
Gradient descent and momentum
A model cannot see the valley it is searching for. It can only feel the slope under its feet, and take a step.
New ideas gradientgradient descentlearning ratelocal minimummomentumadaptive steps (Adam)feature scaling
Episode 3 · London, 1958
“By how surprised I am.”
Logistic regression and classification
Two crowds of points and a line between them. Distance from the line becomes a probability, and every point that surprises the model pulls the line toward it.
New ideas classificationdot productweight vectorsigmoidpredicted probabilitydecision boundarycross-entropylikelihoodaccuracy
Episode 4 · Berkeley, 1951
“By what my neighbours say.”
k-nearest neighbours, distance and scaling
Two hundred and forty houses on a turquoise town plan. Every spot on the map asks its k nearest houses and takes their vote, and turning k wears the town down from jagged islands to smooth hills, then sinks it to sea level.
New ideas distancenearest neighboursk-nearest neighboursmodels without a fixed shape (non-parametric)curse of dimensionality
Act two
A score on the data you fitted is not the truth, and one number is never the whole story: which mistakes, and for whom.
The TradeThe Honest ExamThe Rare CaseWrong for Whom?
Episode 5 · Providence, 1992
“Wrong the same way, or wrong every time.”
The bias–variance trade-off
Two thousand models, each trained on its own thirty noisy points. Too simple and they are wrong the same way. Too flexible and they are wrong every time.
New ideas generalisationtraining errortest erroroverfittingunderfittingbiasvarianceirreducible noisemodel complexitypolynomial features
Episode 6 · London, 1974
“On what I haven’t seen.”
Regularisation and cross-validation
A penalty tames a model that bends too much. Folds of held-out data choose how hard to press it. One peek at the answers turns fifty coin flips into a 97% score.
New ideas regularisationridge (L2)lasso (L1)sparsityhyperparametertrain / test splitvalidation setcross-validationdata leakagesplitting by time
Episode 7 · London, 1763
“It depends which way I’m wrong.”
Confusion matrix, precision and recall, ROC, base rates and calibration
A made-up town of 100,000 takes a test that catches 99% of the ill. Everyone who tests positive walks into one queue, and only one in six of them is ill. Then a real thyroid model, its threshold, its curves and its honesty.
New ideas confusion matrixfalse positivefalse negativeprecisionrecalldecision thresholdROC curvearea under the curve (AUC)base rateBayes' ruleclass imbalancecalibration
Episode 8 · Ithaca, New York, 2016
“More for some of you than others.”
Fairness: error rates by group and the impossibility
Two invented groups, one risk model, one threshold, and more than twice the false alarms for one group. Give each group its own threshold and try every pair: with different base rates, a flag cannot mean the same and the mistakes fall equally at once, so fair is a choice someone has to state.
New ideas error rates by groupequal selection rates (demographic parity)equal error rates (equalised odds)the fairness impossibilityproxy features
Part two · episodes 9–18
Straight lines are not enough. It tries lifting, asking and correcting; then it loses its teacher; then it learns to bend space itself.
Act three
Split the data with questions, correct the last mistake, then ask the model why it answered.
Two Hundred TreesThe Last MistakeWhy Did It Say That?
Episode 9 · Berkeley, 2001
“Wrong alone, right together.”
Decision trees and random forests
One tree remembers every point it studied and stumbles on new ones. Two hundred trees, each a little wrong, agree on something better.
New ideas decision treeimpuritybootstrapensemblemajority votedecorrelated errors
Episode 10 · Stanford, California, 2001
“By what’s left over.”
Gradient boosting
Boosting grows trees in a line, each one fitted to what the others still get wrong: gradient descent where every step is a tree. Watch Downhill's landscape carved out of flat steps, then see it lose, honestly, to a forest on this terrain.
New ideas boostingweak learnerfitting residualsshrinkageearly stopping
Episode 11 · Santa Monica, 1951
“Shared out among my reasons.”
Feature importance, partial dependence and Shapley values
A boosted model prices 20,640 California districts from the 1990 census. Shuffle its inputs, turn one dial for everyone, and share one answer out fairly among its reasons. At dusk the coast lights up with the model's location premium.
New ideas feature importancepermutation importancepartial dependenceShapley valueslocal explanation
Act four
No answers given. Find the groups, the soft edges and the directions that matter.
Five PinsSoft EdgesShadows
Episode 12 · Murray Hill, New Jersey, 1957
“By the distance to my pin.”
k-means clustering
Four hundred thousand stars that nobody has named. k-means finds the clumps without being told what a clump is.
New ideas unsupervised learningclusteringcentroidinertiak-meansinitialisationelbow method
Episode 13 · Cambridge, Massachusetts, 1977
“By how unlikely my story is.”
Gaussian mixtures and EM
The same 400,000 stars as Five Pins, but now a star may belong a little to two clumps. Starting from Five Pins' own answer, EM runs live in your browser, ellipses stretch and turn round by round, and the long bar k-means cut in two ends up inside one ink ring.
New ideas Gaussiancovariancemixture modelsoft assignmentexpectation–maximisation
Episode 14 · London, 1901
“By what the shadow loses.”
Principal component analysis
Turn a light around a cloud of points until its shadow is as wide as it can be. That direction keeps the most of the story, and a few dozen such directions stand in for all 256 pixels of a real handwritten digit.
New ideas projectionprincipal componentexplained variancedimensionality reductioneigenvectorreconstruction
Act five
A lift chosen by hand, then lifts the model learns for itself: many small sums, wired together and bent, until they read.
The Widest StreetFolding SpaceThe City That ReadsThe Sliding Window
Episode 15 · Holmdel, New Jersey, 1992
“By standing too close.”
Support vector machines and kernels
Hundreds of lines split two crowds perfectly. The support vector machine picks the one with the widest empty street, set by the few points on its kerbs, and a kernel lifts the data until a flat street can curve.
New ideas marginsupport vectorshinge losssoft marginkernelfeature mapkernel trick
Episode 16 · Urbana, Illinois, 1989
“Wrong until I bend.”
Hidden layers and nonlinearity
Two spirals no straight line can split. A small network stretches, turns and bends the plane, layer by layer, until a flat cut separates them, and the cut, carried back, is a spiral.
New ideas neuronactivation functionhidden layernonlinearitylearned representationlinear separability
Episode 17 · San Diego, 1986
“Every road shares the blame.”
Neural networks and backpropagation
Each tower is a neuron, each road a weight, each lit window a number it is carrying. The city teaches itself to read digits while you watch.
New ideas backpropagationchain rulesoftmaxmulti-class outputminibatchepoch
Episode 18 · Holmdel, New Jersey, 1989
“The moment it moves.”
Convolutional networks
Move a real handwritten digit two pixels and a dense network forgets it. A convolution slides one small filter across the whole image and asks only whether a shape appeared, not where, so the network still recognises a digit after it moves sideways. Up and down it slips, and the film shows why.
New ideas convolutionfilteractivation mapweight sharingpoolingreceptive fieldshift tolerance
Part three · episodes 19–25
Words become places, words watch each other, and at last it speaks. Then it acts, and then it makes.
Act six
Words become numbers, words look at each other, and a model writes one token at a time.
A Map of WordsEvery Word WatchesThe Next Word
Episode 19 · Mountain View, 2013
“About the company a word keeps.”
Word embeddings
A word can be a list of numbers learned from the company it keeps. Your browser learns them for the Alice books, then you fly over a map of 10,000 words where closeness is a dot product.
New ideas tokenembeddingcosine similaritycontext windowdirections as relations
Episode 20 · Mountain View, 2017
“About who to listen to.”
Self-attention in transformers
Every word watches every other word. Real weights from a real language model, BERT, reading one sentence, drawn as a star chart.
New ideas attentionquery, key, valueattention weightsscaled dot productmultiple headscontextual meaningword order (positional encoding)
Episode 21 · Murray Hill, New Jersey, 1948
“By the surprise of the next word.”
Language models
A real GPT-2 reads "The animal didn't cross the street because it was too" and lights all 50,257 tokens it could say next. Writing is picking one, adding it, and asking again.
New ideas language modelnext-token predictionlogitstemperaturesamplingautoregressioncausal maskperplexity
Act seven
Learning to act from rewards, learning what people want, telling cause from coincidence, and creating from noise.
The Maze That Learns BackTaught to HelpDid It Work?From Noise
Episode 22 · Cambridge, England, 1989
“About the future, a little less each time.”
Reinforcement learning (Q-learning)
No answers, only rewards. An agent wanders a maze, and the value of every square rises backward from the goal until the arrows agree and the path straightens.
New ideas agentstateactionrewardepisodeQ-valueBellman updatediscountexplorationpolicyvalue function
Episode 23 · San Francisco, 2019
“By what you would have preferred.”
From next-word model to assistant: fine-tuning and RLHF
The same question goes to two real language models: SmolLM2, trained only to continue text, and its instruction-tuned twin. A reward model learned from people's choices then pulls 64 of the base model's replies on a leash called β, until the light pours into confident wrong answers.
New ideas pre-trainingfine-tuninginstruction tuninghuman preference comparisonsreward modelreinforcement learning from human feedbackKL penalty
Episode 24 · Harpenden, England, 1925
“Not until somebody changes something.”
Experiments and causation: randomisation, confounding, Simpson’s paradox
In 1973 Berkeley admitted men at a higher rate across the whole campus, yet women at a higher rate in four of its six largest departments. In an invented shop, keen customers fake a +7.93-point win. A coin flip finds +0.77, and 1,000 resamples say how unsure that is.
New ideas correlation versus causationconfounderrandomisationA/B testSimpson's paradoxconfidence intervalselection bias
Episode 25 · Berkeley, 2020
“About the noise, one step at a time.”
Diffusion models
Dissolve Chapter 1's world into static, then teach a small network to undo one sliver of noise at a time. Run it from pure static and 100,000 new points draw that world again, and Chapter 1's line fits itself to them.
New ideas generative modelforward noisingnoise predictionreverse processnoise schedule