Notebook 7 of 9
Section 6: TurboQuant — Rotating Your Way to Better Cache Compression
Let us think carefully about what quantization actually does. When you quantize a vector — a list of numbers — you are replacing each number with the nearest value on a fixed grid. If you have 4-bit quantization, that grid has 16 levels. Those 16 levels must span the full range of the vector from its smallest to its largest value. Everything in between gets rounded to the nearest grid line.
Ready to Code
Download this notebook and open it in Google Colab. Work through the exercises — this notebook includes voice narration inside Colab.
~25 min4 exercises
0/9 complete