D8 · Mechanistic Interpretability
Sparse Autoencoders
Describe how an SAE is trained and what a learned feature dictionary contains.
This is the whole tool. Take a layer's activations, spread them into a much wider layer, then squeeze them back. One rule makes it useful: the wide middle must stay nearly empty.