Attribution Methods: Saliency, Integrated Gradients, LIME, and SHAP
What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.
What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.
How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.
Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.
Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.
The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.
Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.
Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.
What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.
How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.
Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.
Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.
The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.
Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.
Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.
A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.
A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.