Stage 3 · Expert · D8
Mechanistic Interpretability
The dark matter of AI.
11 lessons · 126 minResearch
About this chapter
We can read every weight in a model and still not know what it is doing. Mechanistic interpretability tries to close that gap by finding features, then circuits, then a story about computation you can test. This chapter covers superposition, the reason a neuron can mean five things at once, and sparse autoencoders, the current best tool for pulling those meanings apart. You will patch activations to prove a claim rather than assert it, and you will see plainly where the field still fails.
What you will be able to do
- 110 min
What Counts as a Feature
Define a feature as a direction in activation space, not a neuron.
- 211 min
Neurons, Features, Circuits
Trace a small circuit that combines features into a decision.
- 310 min
The Neuron That Means Five Things
Recognize polysemantic neurons and say why they are the normal case.
- 412 min
Superposition
Explain how a model packs more features than it has dimensions.
- 513 min
Sparse Autoencoders
Describe how an SAE is trained and what a learned feature dictionary contains.
- 611 min
Probing
Train a probe and say why a probe finding something is weak evidence of use.
- 713 min
Activation Patching
Design a patching experiment that isolates one component's causal role.
- 812 min
Induction Heads
Explain the copy-and-continue pattern and why it matters for in-context learning.
- 911 min
Attribution Maps and Their Failures
Read a saliency map and describe a test that shows it can be meaningless.
- 1011 min
What We Cannot Read Yet
State the current limits of interpretability honestly, with examples.
- 1112 min
Why Safety Needs This
Connect interpretability results to concrete safety decisions.