Series
Mechanistic Interpretability
How do transformers actually work inside? From representation geometry and attention circuits to sparse autoencoders, steering vectors, and attribution methods — a rigorous series on understanding what neural networks compute and why.
- 1
Representation Geometry: How Neural Networks Encode Meaning
The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.
2025-06-20
- 1
Representation Geometry: How Neural Networks Encode Meaning
The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.
2025-06-20
- 2
What Each Transformer Component Actually Does
Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.
2025-06-20
- 2
What Each Transformer Component Actually Does
Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.
2025-06-20
- 3
Logit Lens: How Predictions Form Layer by Layer
Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.
2025-06-20
- 3
Logit Lens: How Predictions Form Layer by Layer
Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.
2025-06-20
- 4
Sparse Autoencoders: Decomposing Neural Networks into Interpretable Features
Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.
2025-06-20
- 4
Sparse Autoencoders: Decomposing Neural Networks into Interpretable Features
Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.
2025-06-20
- 5
Circuits: How Transformers Implement Algorithms
How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.
2025-06-20
- 5
Circuits: How Transformers Implement Algorithms
How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.
2025-06-20
- 6
Activation Steering: A New Frontier in AI Control—But Does It Scale?
A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.
2025-02-02
- 6
Activation Steering: A New Frontier in AI Control—But Does It Scale?
A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.
2025-02-02
- 7
Attribution Methods: Saliency, Integrated Gradients, LIME, and SHAP
What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.
2025-06-20
- 7
Attribution Methods: Saliency, Integrated Gradients, LIME, and SHAP
What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.
2025-06-20
- 8
Open Problems in Mechanistic Interpretability
Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.
2025-06-20
- 8
Open Problems in Mechanistic Interpretability
Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.
2025-06-20