The dark matter of AI.
D8 · Mechanistic Interpretability
We can read every weight in a model and still not know what it is doing. Mechanistic interpretability tries to close that gap by finding features, then circuits, then a story about computation you can test. This chapter covers superposition, the reason a neuron can mean five things at once, and sparse autoencoders, the current best tool for pulling those meanings apart. You will patch activations to prove a claim rather than assert it, and you will see plainly where the field still fails.
- What Counts as a Feature
- Neurons, Features, Circuits
- The Neuron That Means Five Things
- Superposition
- Sparse Autoencoders
- Probing
- Activation Patching
- Induction Heads
- Attribution Maps and Their Failures
- What We Cannot Read Yet
- Why Safety Needs This