VizuaraVizuara AI Pods

MXFP4 Quantization-Aware Training from Scratch

Block-scaled 4-bit floating point (MXFP4) quantization-aware training: why frontier open models like Kimi K3 ship 4-bit weights, how block scaling works (MXFP4 weights + MXFP8 activations, per-block scale factors), and why learning to compensate for quantization error during SFT beats post-hoc PTQ. From-scratch build: implement block-scaled quantization, a fake-quant QAT training loop, and compare PTQ vs QAT on a small LM.

intermediate~5 hours9 notebooks

Curator of this Module

RD

Raj Dandekar

Checking access…

Learning Path

Article
1
Introduction
2
Notebook 2
3
Notebook 3
4
Notebook 4
5
Notebook 5
6
Notebook 6
7
Notebook 7
8
Notebook 8
9
Conclusion
Case Study
Certificate