Kiru Lab  /  Track

Mechanistic Interpretability

To cut. The point where the curriculum becomes the rest of Kiru.

Mechanistic interpretability tries to recover the algorithm a network implements — not a story about its behavior, but a description of the computation that survives intervention. It is an inverse problem, and the discipline it demands is the one this whole curriculum has been building toward: state a hypothesis, design the intervention that would falsify it, run it across seeds, and report what happened. This is the track where a DEM-X entry stops being a description and becomes a diagnosis.

The question this track answers

Can you state what a model computed, and prove it by intervention rather than by narration?

By the end you can

  • Distinguish a correlational observation from a causal claim about a model
  • Run activation patching and interpret the result correctly
  • Design an ablation that isolates a component's contribution
  • Explain superposition and why it makes single neurons misleading
  • Turn a mechanistic finding into a reproducible DEM-X submission

Come here from

Scope

7 lessons, roughly 26 hours.