Kiru Lab  /  Mechanistic Interpretability

Features and Representation

What a model represents, why neurons rarely mean one thing, and what a circuit is.

In this module you will

  • Define a feature as a direction rather than a neuron
  • Explain superposition and polysemanticity
  • Describe a circuit as components composing across layers