Kiru Lab  /  Language Models

Anatomy of a Transformer

Attention, the residual stream, and the block that repeats until you have a model.

In this module you will

  • Implement scaled dot-product attention
  • Explain the residual stream as a communication bus
  • Describe what MLP layers contribute that attention does not