Linear Attention and the Delta Rule: Building Kimi Delta Attention
From softmax attention to linear attention to the delta rule, building up to Kimi Delta Attention (KDA): per-channel learnable forgetting via a diagonal-plus-low-rank transition, why it powers the 1M-token context of Kimi K3, implemented from first principles with a working long-context toy model.
intermediate~5 hours9 notebooks
Curator of this Module
RD
Rajat Dandekar
Checking access…
Learning Path
Article
1
Introduction2
Notebook 23
Notebook 34
Notebook 45
Notebook 56
Notebook 67
Notebook 78
Notebook 89
ConclusionCase Study
Certificate