MXFP4 Quantization-Aware Training from Scratch
Block-scaled 4-bit floating point (MXFP4) quantization-aware training: why frontier open models like Kimi K3 ship 4-bit weights, how block scaling works (MXFP4 weights + MXFP8 activations, per-block scale factors), and why learning to compensate for quantization error during SFT beats post-hoc PTQ. From-scratch build: implement block-scaled quantization, a fake-quant QAT training loop, and compare PTQ vs QAT on a small LM.
intermediate~5 hours9 notebooks
Curator of this Module
RD
Raj Dandekar
Checking access…
Learning Path
Article
1
Introduction2
Notebook 23
Notebook 34
Notebook 45
Notebook 56
Notebook 67
Notebook 78
Notebook 89
ConclusionCase Study
Certificate