43 chapters · 550 lessons · about 92 hours
From one neuron to agents
One road in three stages. The level test finds where you join it.
Stage 1 · Beginner
Spark
How a machine learns, from one neuron up.
No background needed. By the end you can explain, and train, a small neural network.
F0
Start Here
Your phone is full of models that learned from examples. Here is what they are.
10 lessons · 97 minGentle
D1
The Perceptron
A model like ChatGPT is built from millions of these.
12 lessons · 141 minEasy
F7
Where AI Came From
The first AI researchers asked for one summer. The work is now seventy years old.
15 lessons · 121 minGentle
- 1What People Mean by AI7′
- 2Rules Written by Hand8′
- 3Thinking Ahead9′
- 4The Shortest Way There8′
- 5Experts in a Box8′
- 6The Knowledge Bottleneck8′
- 7Two Winters8′
- 8Learning from Examples Instead8′
- 9Games as Milestones9′
- 10Data and Chips8′
- 11Attention and Scale8′
- 12Chat Goes Public7′
- 13AI in the Arab World9′
- 14What It Can and Cannot Do8′
- 15Reading AI News Without the Hype8′
F1
Numbers That Learn
A photo is a single point in a space with three million dimensions.
12 lessons · 105 minGentle
F2
The Slope of Everything
Every AI model is a machine for taking one derivative.
12 lessons · 101 minGentle
- 1The Question Every Model Asks7′
- 2How Fast, Not How Much7′
- 3Zoom In Until It Is Straight8′
- 4The Derivative Is a Slope at a Point9′
- 5The Rules You Will Actually Use8′
- 6The Chain Rule Is the Whole Game10′
- 7One Knob at a Time8′
- 8The Gradient Points Uphill9′
- 9Bottoms, Tops and Saddles9′
- 10Every Model Is a Derivative Machine9′
- 11Measure It Both Ways9′
- 12A Taste of Area8′
D2
Gradient Descent
The fear that held neural networks back was aimed at the wrong thing.
14 lessons · 176 minEasy
- 1Loss Is a Landscape10′
- 2The Slope Tells You Which Way Is Down11′
- 3One Step Downhill12′
- 4Too Big, Too Small12′
- 5Bowls and Real Landscapes12′
- 6The Wrong Fear15′
- 7Noisy Steps Beat Perfect Ones15′
- 8Give the Ball Some Weight12′
- 9A Different Step for Every Weight12′
- 10Adam: Both Ideas at Once15′
- 11Change the Step as You Go13′
- 12The Spike That Ruins a Week12′
- 13Knowing When to Stop13′
- 14Watch Them Race12′
M1
The Best Straight Line
One equation, published in 1805, still runs inside every neural network.
12 lessons · 128 minEasy
- 1Fit It by Eye First8′
- 2Why We Square the Error10′
- 3The Loss Surface Is a Bowl10′
- 4The Exact Answer11′
- 5The Same Answer, Step by Step11′
- 6More Than One Input11′
- 7Put the Features on One Scale10′
- 8Curves From a Linear Model11′
- 9Trying Too Hard12′
- 10Ridge and Lasso as a Budget12′
- 11R Squared and Its Traps11′
- 12Where This Lives Inside Every Network11′
F3
Being Wrong on Purpose
A chat model never picks a word. It rolls a weighted die.
13 lessons · 126 minEasy
- 1A Number You Do Not Know Yet7′
- 2Coins and Categories9′
- 3The Bell and the Flat Line10′
- 4Where It Sits and How Much It Wobbles9′
- 5Joint, Marginal, Conditional10′
- 6Bayes Is a Belief Update12′
- 7When Knowing One Tells You Nothing7′
- 8Sampling and the Law of Large Numbers10′
- 9Which Setting Explains the Data Best10′
- 10Surprise, Measured10′
- 11Cross-Entropy: The Loss You Will Use Forever11′
- 12KL Divergence10′
- 13Softmax and Temperature11′
M2
Drawing the Line
Spam filters, credit scores and cancer screens are all one question: which side of the line?
14 lessons · 131 minEasy
- 1From a Line to a Label8′
- 2The Soft Switch8′
- 3Logistic Regression10′
- 4The Loss That Punishes Confidence9′
- 5Where the Numbers Came From10′
- 6Boundaries You Can See12′
- 7The Model That Never Trains9′
- 8The Four Boxes9′
- 9Caught and Cried Wolf10′
- 10Choosing a Cut From Money10′
- 11Every Threshold at Once10′
- 12When the Answer Is Almost Always No9′
- 13More Than Two Answers9′
- 14The Last Layer of Everything8′
D3
Backpropagation
The F = ma of artificial intelligence.
14 lessons · 154 minSteady
- 1Every Model Is a Graph11′
- 2The Forward Pass10′
- 3Every Node Knows One Small Derivative13′
- 4Chaining Backwards13′
- 5One Neuron, By Hand11′
- 6Two Layers, By Hand12′
- 7Backprop in Matrices11′
- 8Reverse-Mode Autodiff10′
- 9Why It Costs So Little9′
- 10Gradients That Fade10′
- 11Gradients That Explode, Units That Die10′
- 12Autograd, For Real12′
- 13Checking Gradients Numerically11′
- 14What Backprop Is Not11′
D4
Deep Learning
The geometry of depth.
14 lessons · 164 minSteady
- 1Layers Fold Space12′
- 2Folds Made of ReLU11′
- 3One Layer Can Do Anything11′
- 4Depth Is Cheaper Than Width12′
- 5Features Nobody Designed11′
- 6Where the Weights Start Matters10′
- 7Batch Norm and Layer Norm12′
- 8Dropout and Weight Decay12′
- 9The Shortcut That Opened Up Depth12′
- 10Five Lines, In Order13′
- 11Memorizing Versus Understanding12′
- 12Double Descent11′
- 13Build an MLP on Moons13′
- 14Reading the Curves12′
Stage 2 · Intermediate
Craft
Networks that see, remember and write.
You learn the tools of the trade and open up the transformer. By the end you know what happens inside a chat model.
D5
Seeing
The moment we stopped understanding our own models.
13 lessons · 147 minSteady
- 1An Image Is a Tensor9′
- 2One Filter, Dragged Across12′
- 3Kernels You Can Draw11′
- 4Padding and Stride11′
- 5Pooling10′
- 6How Much Can This Neuron See?12′
- 7Channels and Depth12′
- 8What Actually Won in 201212′
- 9Looking Inside a Vision Model12′
- 10Going Deep With ResNet11′
- 11Reusing a Trained Model12′
- 12When Attention Came for Vision11′
- 13A Panda, One Nudge, a Gibbon12′
F5
Speaking to Machines
Under the hood, a model is mostly arithmetic on grids of numbers, and you can learn to write it.
13 lessons · 111 minGentle
D6
Sequences & Memory
Before attention, a model had to squeeze a whole paragraph into one vector.
11 lessons · 107 minDemanding
- 1Data Where Order Matters8′
- 2A Vector That Remembers9′
- 3Unrolling the Loop10′
- 4Backprop Through Time11′
- 5The Same Problem, Now in Time10′
- 6Gates That Choose What to Keep12′
- 7GRU: Fewer Gates, Similar Result8′
- 8Training vs Generating9′
- 9A Character-Level Language Model11′
- 10One Vector for a Whole Sentence10′
- 11The Case for Attention9′
F4
Reading Data Honestly
Train the same model twice with different random seeds and you get two scores. Some published wins are smaller than that gap.
12 lessons · 114 minEasy
- 1Nobody Is Average7′
- 2Shapes Real Data Takes9′
- 3Two Columns That Move Together9′
- 4Who Is Missing From the Sample9′
- 5Luck Has a Size9′
- 6The Interval Is the Result9′
- 7Pull the Test Set Up by Itself10′
- 8Could Luck Have Done It11′
- 9Right Nineteen Times in Twenty10′
- 10Keep Trying and You Will Win10′
- 11Decide the Size Before You Look10′
- 12Reading a Table With Suspicion11′
M7
Predicting Tomorrow
In a forecasting contest of 5,507 teams, most could not beat the rule that tomorrow looks like today.
15 lessons · 117 minDemanding
- 1Yesterday Is a Clue6′
- 2Three Things in One Line7′
- 3The Clocks Inside a Series7′
- 4Tomorrow Looks Like Today8′
- 5Keeping Score7′
- 6Average the Recent Past8′
- 7Let Old Days Fade8′
- 8The Series Predicts Itself9′
- 9Two Calendars, One Shop9′
- 10No Peeking8′
- 11Walk Forward9′
- 12A Range, Not a Number8′
- 13Big Models Against the Baseline8′
- 14Surprises8′
- 15When Not to Forecast7′
L1
Words as Numbers
The word strawberry has three r's, and a model cannot see any of them.
12 lessons · 117 minSteady
- 1Text Has to Become Numbers8′
- 2Characters, Words, Subwords9′
- 3Byte Pair Encoding, Merge by Merge11′
- 4How Big Should the Vocabulary Be?9′
- 5Special Tokens and Chat Templates8′
- 6Embeddings Are Learned Coordinates11′
- 7Similarity and Analogies10′
- 8Meaning From Company10′
- 9Where a Token Sits10′
- 10The Weird Failures11′
- 11What Arabic Costs a Tokenizer12′
- 12Counting Tokens and Money8′
L2
Attention
One lab cut a model's memory of a conversation by 93 percent and kept every head.
11 lessons · 115 minSteady
L3
The Transformer
One architecture ate the whole field.
13 lessons · 134 minSteady
- 1Anatomy of One Block10′
- 2The Feed-Forward Layer10′
- 3Norms and Residuals10′
- 4The Residual Stream10′
- 5Positions in Practice10′
- 6Three Shapes of Transformer11′
- 7GPT and BERT10′
- 8Where Facts Live10′
- 9Mixture of Experts11′
- 10From Vector to Next Token9′
- 11Why It Beat the RNN10′
- 12Reading a Real Config11′
- 13Build a Tiny Transformer12′
F6
Data Is the Model
The network that started the deep learning boom learned from 1.2 million photos, each one labeled by a person.
12 lessons · 118 minEasy
- 1Rows, Features, Labels9′
- 2Pictures and Words as Rows9′
- 3Where Labels Come From10′
- 4Cleaning Without Lying10′
- 5The Holes in the Table10′
- 6Three Piles10′
- 7The Answer That Slipped In10′
- 8When Tomorrow Is Different11′
- 9The One Case Accuracy Hides10′
- 10More Data Without New Facts9′
- 11Who Is in the Table?10′
- 12The Card That Ships With the Data10′
M3
Twenty Questions
The model that wins most tabular competitions is not a neural network.
12 lessons · 113 minSteady
- 1A Tree of Questions8′
- 2Which Question to Ask First10′
- 3Why Tree Boundaries Are Staircases8′
- 4A Tree That Memorizes9′
- 5Pruning and Depth Limits9′
- 6Many Trees, Many Samples9′
- 7Random Forests10′
- 8Learning From Your Own Mistakes10′
- 9Gradient Boosting11′
- 10Feature Importance and Its Traps9′
- 11Crediting Each Feature Fairly10′
- 12When Trees Beat Deep Learning10′
M4
Finding Groups Nobody Labeled
The model has no answers to copy. It has to invent the categories.
12 lessons · 112 minSteady
- 1Near Means Similar8′
- 2Who Is Nearest8′
- 3k-Means, Step by Step10′
- 4Where k-Means Breaks11′
- 5How Many Groups Are There?8′
- 6Clusters Inside Clusters9′
- 7Density, Not Centers9′
- 8PCA Is a Rotation11′
- 9Maps of High-Dimensional Data10′
- 10Points That Belong Nowhere9′
- 11A Grid With Less in It10′
- 12The Curse of Dimensionality9′
M5
Did It Actually Work?
A model that scores 99% can be worse than a coin flip.
13 lessons · 125 minSteady
- 1The Score That Lies8′
- 2Beat the Dumbest Thing First8′
- 3One Metric Per Task11′
- 4Every Example Gets a Turn10′
- 5Splits That Respect Reality9′
- 6Too Stiff or Too Eager10′
- 7What the Curve Tells You to Do10′
- 8Searching Without Fooling Yourself11′
- 9The Wobble Between Runs9′
- 10Look at the Mistakes10′
- 11Hunting Data Leakage11′
- 12Benchmarks Wear Out9′
- 13Reporting Results Honestly9′
M6
What to Show Next
About 80% of the hours people stream on Netflix start with a pick the app made, not a search.
14 lessons · 131 minSteady
- 1Too Much to Show8′
- 2Near Tastes9′
- 3People Like You10′
- 4Two Thin Tables9′
- 5Fitting Only What You Know11′
- 6A Map of Taste9′
- 7The Dish Nobody Has Rated9′
- 8A Click Is Not a Like9′
- 9Find, Then Rank9′
- 10Scoring a List10′
- 11Offline Wins, Online Losses9′
- 12The Feed That Feeds Itself10′
- 13Varied, New and Fair10′
- 14One Space for Search and Suggestions9′
D7
Neural Scaling Laws
AI cannot cross this line, and nobody knows why the line is there.
11 lessons · 112 minDemanding
- 1Reading a Log-Log Plot10′
- 2Power Laws Everywhere10′
- 3Compute, Data, Parameters10′
- 4The Kaplan Result10′
- 5Chinchilla Changes the Recipe11′
- 6Spending a Compute Budget10′
- 7Emergence and Its Critics11′
- 8Why Bigger Keeps Working10′
- 9What a Frontier Run Costs10′
- 10Scaling at Inference Time10′
- 11What the Curve Does Not Predict10′
D10
Learning by Trying
AlphaGo Zero never saw a human game of Go. It beat the version that had, 100 games to 0.
15 lessons · 138 minResearch
- 1Nobody Gives It the Answer8′
- 2Stay or Try Something New8′
- 3Mostly Greedy, Sometimes Curious9′
- 4Where You Stand Changes the Choice9′
- 5A Riyal Today9′
- 6How Good Is This Square?9′
- 7Learning One Step at a Time11′
- 8Curious Early, Careful Late9′
- 9Too Many Squares for a Table11′
- 10Push Up What Worked10′
- 11A Coach Beside the Player10′
- 12It Learned the Score, Not the Task8′
- 13Playing Against Yourself10′
- 14From Games to Robots9′
- 15The Same Loop, Pointed at Words8′
L4
Training a Language Model
Pretraining never hands the model a single fact. It picks facts up because guessing the next word well is impossible without them.
14 lessons · 140 minDemanding
- 1Guess the Next Token9′
- 2Grading the Guess10′
- 3Where the Text Comes From10′
- 4Cleaning the Text10′
- 5Windows, Batches, Packing9′
- 6The Loop at Scale10′
- 7Warmup, Then Decay10′
- 8Small Numbers, Big Spikes10′
- 9Splitting the Work10′
- 10Tokens Against Parameters11′
- 11From Predictor to Assistant10′
- 12Learning From Preferences11′
- 13DPO, and Doing Less10′
- 14Reading a Benchmark Table10′
L5
Talking to Models
In a 2023 test, GPT-3.5 did worse with the answer buried in the middle of twenty documents than with no documents at all.
13 lessons · 136 minSteady
- 1Where a Prompt Goes9′
- 2Three Voices, One Document10′
- 3Say the Thing You Want10′
- 4Show, Do Not Tell11′
- 5Room to Think11′
- 6One Dial Called Temperature11′
- 7Cutting the Tail10′
- 8Output You Can Parse10′
- 9Lost in the Middle10′
- 10Why Models Make Things Up11′
- 11Prompt Injection Is a Security Bug11′
- 12Test Sets, Not Vibes11′
- 13When Prompting Is the Wrong Tool11′
Stage 3 · Expert
Frontier
Build real systems, then read the edge of the field.
You ship with models, train your own, and rebuild landmark papers. By the end you can pick an open problem.
L6
Grounding Models
The model does not know your data. Give it eyes.
13 lessons · 124 minSteady
- 1Why the Model Does Not Know Your Data8′
- 2Search by Meaning10′
- 3Chunking Decides Everything11′
- 4Vector Indexes10′
- 5Keywords Still Matter10′
- 6Read the Shortlist Properly9′
- 7Answers With Receipts9′
- 8Retrieval in Arabic10′
- 9Measuring the Two Halves11′
- 10Four Ways It Fails10′
- 11Paste It All, or Go and Find It9′
- 12Teach It, or Show It9′
- 13Every Tool Is a Retriever8′
L7
Agents
A model in a loop with tools is a different kind of software.
12 lessons · 119 minDemanding
- 1The Loop9′
- 2Tools Are Functions With Descriptions10′
- 3Tool Output Is Data, Not Orders10′
- 4Planning10′
- 5Memory Between Steps10′
- 6More Than One Agent10′
- 7Tool Protocols9′
- 8Guardrails and Permission Models11′
- 9Loops, Hallucinated Tools and Stalls10′
- 10Evaluating Agents11′
- 11Keeping the Bill Sane9′
- 12What They Can Do Today10′
L9
Models That Think Longer
On many problems, letting a model think longer beats making it fourteen times bigger.
14 lessons · 135 minResearch
- 1One Guess Is Not Enough8′
- 2Steps on the Page9′
- 3Ask Again, Take the Vote9′
- 4Thinking Time or a Bigger Model10′
- 5A Checker Picks the Best10′
- 6Grade the Answer or the Working10′
- 7Rewards You Can Check11′
- 8What the Training Grew10′
- 9The Reasoning Models9′
- 10Searching a Tree of Thoughts10′
- 11Tools in the Middle of a Thought9′
- 12Teaching Small Models to Think9′
- 13When Thinking Goes Wrong10′
- 14Measuring Reasoning Honestly11′
A1
Building With Model APIs
Most AI products are two hundred lines of glue around one API call.
11 lessons · 107 minSteady
A3
Shipping and Operating AI
The model is ten percent of the system.
12 lessons · 116 minDemanding
- 1The Model Is Ten Percent8′
- 2Two Phases and a Cache10′
- 3Batching and the Tail11′
- 4The Cost of One Answer10′
- 5Watching Quality, Not Uptime10′
- 6When the World Moves10′
- 7Loops That Help, Loops That Rot10′
- 8Guardrails and Fallbacks10′
- 9Shipping a Change10′
- 10Versions You Can Roll Back9′
- 11What You Keep9′
- 12When Not to Use AI9′
A2
Training Your Own Small Model
A model of 449 numbers, trained in under a second, can be all your one task needs.
12 lessons · 124 minSteady
- 1A Task Worth a Small Model9′
- 2A Few Hundred Labels10′
- 3The Number You Have to Beat9′
- 4Training a Tiny Classifier12′
- 5Overfit on Purpose10′
- 6Three Ways to Fix It11′
- 7A Language Model You Can Read12′
- 8Fine-Tune or Start From Scratch10′
- 9LoRA, by Picture10′
- 10Letting the Giant Teach10′
- 11Shrinking It to Ship10′
- 12Judging It Against the Giant11′
L10
Smaller, Faster, Cheaper
Most of the time a chat model spends writing to you, its chip is waiting for memory.
14 lessons · 146 minResearch
- 1Who Gets the Answer9′
- 2The Wall Is Memory10′
- 3One Scale or Many11′
- 4One Number Ruins the Row11′
- 5Round, Then Repair12′
- 6A Teacher With a Vocabulary11′
- 7Cutting Weights11′
- 8One Base, Many Adapters10′
- 9Cheap to Run, Costly to Hold10′
- 10A Small Model Drafts10′
- 11How Far Ahead to Guess10′
- 12The Long Conversation Bill10′
- 13A Model in Your Pocket10′
- 14Choosing With a Measured Eval11′
D9
Image & Video Generation
AI video is Brownian motion played backwards.
15 lessons · 178 minDemanding
- 1Squeeze, Then Rebuild10′
- 2Walking Through Latent Space11′
- 3Variational Autoencoders12′
- 4Two Networks in a Fight12′
- 5Why GANs Collapse12′
- 6Destroying an Image on Purpose12′
- 7Learning to Undo One Step13′
- 8Following the Data Uphill12′
- 9Sampling Schedules11′
- 10Conditioning and Guidance13′
- 11Diffusion in Latent Space12′
- 12From Words to Pictures12′
- 13Making Time Consistent12′
- 14Scoring a Picture Maker12′
- 15What These Models Understand12′
D11
Seeing and Hearing Together
To a transformer, a photo, a voice note and a sentence are the same thing: a line of tokens.
14 lessons · 135 minResearch
- 1Every Sense Becomes a Line9′
- 2Pictures Meet Their Captions8′
- 3Right Pairs Up, Wrong Pairs Down10′
- 4Classify by Writing Captions9′
- 5Pictures Inside a Language Model10′
- 6What a Picture Costs9′
- 7Reading a Page11′
- 8Sound as a Picture10′
- 9From Voice to Words10′
- 10Words Back to Voice10′
- 11Pictures in Time10′
- 12Making, Not Only Reading8′
- 13Where They Fail10′
- 14Models That Act11′
D8
Mechanistic Interpretability
The dark matter of AI.
11 lessons · 126 minDemanding
L8
Safety, Alignment & Society
We built minds we cannot fully inspect. Now what?
15 lessons · 157 minDemanding
- 1What Alignment Means10′
- 2Doing Exactly What You Asked10′
- 3Searching Harder Makes It Worse11′
- 4The Model Learns Your Taste10′
- 5Confident and Wrong10′
- 6Answer or Stay Quiet11′
- 7Bias You Can Measure11′
- 8What the Model Remembers10′
- 9Why a Refusal Breaks11′
- 10Going Looking on Purpose11′
- 11An Evaluation You Can Trust11′
- 12Looking Inside11′
- 13Misuse, Accidents and Uplift10′
- 14The Rules, in Plain Words10′
- 15Electricity, Work and What Nobody Knows10′
A4
Capstones
Six projects that run on this device, no server anywhere.
11 lessons · 126 minDemanding
- 1Scoping a Project You Will Finish10′
- 2Project 1: A Spam Filter12′
- 3Project 1: The Cut Is a Decision11′
- 4Project 2: A Digit Recognizer12′
- 5Project 2: Where It Confuses Itself11′
- 6Project 3: A Tiny Next-Word Model13′
- 7Project 4: An Answerer That Cites12′
- 8Project 4: Saying I Do Not Know11′
- 9Project 5: Catching a Regression12′
- 10Project 6: Redraw One Figure12′
- 11Writing It Up10′
R1
How to Read a Paper
Read the figures first. The abstract is marketing.
12 lessons · 116 minDemanding
- 1Why Read Papers at All7′
- 2Preprints, Venues and Peer Review8′
- 3The Shape of a Paper9′
- 4Abstract, Then Figures, Then Decide11′
- 5Reading the Math11′
- 6Reading the Results Table11′
- 7Spotting a Weak Baseline10′
- 8Following the Thread9′
- 9Reproducing One Claim13′
- 10Spotting Overclaiming9′
- 11Writing a Summary10′
- 12Keeping a Reading List8′
R2
Landmark Papers, Rebuilt
Fourteen papers built the field. You can rebuild the core of each one.
14 lessons · 154 minResearch
- 1Perceptron (1958)11′
- 2Learning Representations by Back-Propagating Errors (1986)12′
- 3LeNet (1998)11′
- 4AlexNet (2012)11′
- 5Word2Vec (2013)10′
- 6Sequence to Sequence (2014)10′
- 7Attention Is All You Need (2017)13′
- 8GPT-2 and GPT-311′
- 9Scaling Laws (2020)10′
- 10Chinchilla (2022)10′
- 11CLIP (2021)11′
- 12Denoising Diffusion (2020)12′
- 13InstructGPT (2022)11′
- 14DeepSeek V2 and V311′
R3
Open Problems
Nobody knows the answers in this chapter. That is the point.
12 lessons · 121 minResearch
- 1What Counts as an Open Problem8′
- 2Where Does Scaling Stop?11′
- 3Running Out of Text10′
- 4Do Models Reason?11′
- 5Learning After Training10′
- 6Evaluation Is Broken10′
- 7Can We Ever Read a Model?10′
- 8Energy, Chips and Limits9′
- 9Alignment as an Open Problem11′
- 10Senses and Bodies10′
- 11Picking a First Research Project11′
- 12Running It Without Fooling Yourself10′