ZeroShotMind

Series

Mechanistic Interpretability

How do transformers actually work inside? From representation geometry and attention circuits to sparse autoencoders, steering vectors, and attribution methods — a rigorous series on understanding what neural networks compute and why.

Defining Frontier
interpretabilitymechanistic-interpretabilitytransformerscircuitsfeatures
  1. 1

    Representation Geometry: How Neural Networks Encode Meaning

    The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.

    2025-06-20

  2. 1

    Representation Geometry: How Neural Networks Encode Meaning

    The linear representation hypothesis, superposition, polysemanticity, and why transformer activations are more structured than they look.

    2025-06-20

  3. 2

    What Each Transformer Component Actually Does

    Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.

    2025-06-20

  4. 2

    What Each Transformer Component Actually Does

    Attention heads as information-routing circuits, MLP layers as key-value memories, and the residual stream as a shared communication bus.

    2025-06-20

  5. 3

    Logit Lens: How Predictions Form Layer by Layer

    Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.

    2025-06-20

  6. 3

    Logit Lens: How Predictions Form Layer by Layer

    Applying the unembedding matrix at intermediate layers to watch how a transformer's prediction evolves — and what direct logit attribution tells us about which components matter.

    2025-06-20

  7. 4

    Sparse Autoencoders: Decomposing Neural Networks into Interpretable Features

    Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.

    2025-06-20

  8. 4

    Sparse Autoencoders: Decomposing Neural Networks into Interpretable Features

    Dictionary learning for neural networks — how sparse autoencoders recover monosemantic features from polysemantic activations, and what Anthropic's scaling monosemanticity work found in Claude.

    2025-06-20

  9. 5

    Circuits: How Transformers Implement Algorithms

    How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.

    2025-06-20

  10. 5

    Circuits: How Transformers Implement Algorithms

    How to identify the minimal subgraph of attention heads and MLP layers that implements a specific behavior — and what we've learned from the indirect object identification circuit in GPT-2.

    2025-06-20

  11. 6

    Activation Steering: A New Frontier in AI Control—But Does It Scale?

    A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.

    2025-02-02

  12. 6

    Activation Steering: A New Frontier in AI Control—But Does It Scale?

    A practical introduction to activation steering, why it is promising for AI control, and where scaling challenges still remain.

    2025-02-02

  13. 7

    Attribution Methods: Saliency, Integrated Gradients, LIME, and SHAP

    What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.

    2025-06-20

  14. 7

    Attribution Methods: Saliency, Integrated Gradients, LIME, and SHAP

    What each attribution method actually computes, where they agree, where they fail, and whether gradient-based and perturbation-based approaches are still relevant for LLMs.

    2025-06-20

  15. 8

    Open Problems in Mechanistic Interpretability

    Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.

    2025-06-20

  16. 8

    Open Problems in Mechanistic Interpretability

    Faithfulness vs. plausibility, scaling to frontier models, the composition problem, automated interpretability, and what it would take to actually understand a large language model.

    2025-06-20