Kiru Lab / Track
Mechanistic Interpretability
To cut. The point where the curriculum becomes the rest of Kiru.
Mechanistic interpretability tries to recover the algorithm a network implements — not a story about its behavior, but a description of the computation that survives intervention. It is an inverse problem, and the discipline it demands is the one this whole curriculum has been building toward: state a hypothesis, design the intervention that would falsify it, run it across seeds, and report what happened. This is the track where a DEM-X entry stops being a description and becomes a diagnosis.
Can you state what a model computed, and prove it by intervention rather than by narration?
By the end you can
- Distinguish a correlational observation from a causal claim about a model
- Run activation patching and interpret the result correctly
- Design an ablation that isolates a component's contribution
- Explain superposition and why it makes single neurons misleading
- Turn a mechanistic finding into a reproducible DEM-X submission
Module 1
Features and Representation
What a model represents, why neurons rarely mean one thing, and what a circuit is.
Module 2
Causal Methods
Patching, ablation, and probing — the difference between watching and cutting.
Module 3
From Mechanism to Diagnosis
The closing module: turning mechanistic findings into DEM-X evidence, and the standard Kiru holds.