D8 · Mechanistic Interpretability
Probing
Train a probe and say why a probe finding something is weak evidence of use.
Here is the cheapest question in interpretability: is some property written down in this layer? Label activations, train a small classifier to recover the label, and see if it works.