VizuaraVizuara AI Pods

Linear Attention and the Delta Rule: Building Kimi Delta Attention

From softmax attention to linear attention to the delta rule, building up to Kimi Delta Attention (KDA): per-channel learnable forgetting via a diagonal-plus-low-rank transition, why it powers the 1M-token context of Kimi K3, implemented from first principles with a working long-context toy model.

intermediate~5 hours9 notebooks

Curator of this Module

RD

Rajat Dandekar

Checking access…

Learning Path

Article
1
Introduction
2
Notebook 2
3
Notebook 3
4
Notebook 4
5
Notebook 5
6
Notebook 6
7
Notebook 7
8
Notebook 8
9
Conclusion
Case Study
Certificate