VizuaraVizuara AI Pods

ZeRO-1, 2, 3 — Zero Redundancy Optimizer

The elegant idea behind ZeRO: instead of replicating everything on every GPU, partition optimizer states, gradients, and parameters across GPUs.

beginner~4 hours3 notebooksThe Memory Redundancy Problem in Data ParallelismZeRO Stage 1: Partitioning Optimizer StatesZeRO Stage 2: Adding Gradient PartitioningZeRO Stage 3 (FSDP): Partitioning Everything

Curator of this Module

Dr. Rajat Dandekar

Dr. Rajat Dandekar

Course Instructor

Dr. Rajat Dandekar is a researcher and educator specializing in AI/ML, with a passion for making complex concepts accessible through intuitive explanations and hands-on learning.

Checking access…

Learning Path

Article
1
Notebook 1
2
Notebook 2
3
Notebook 3
Certificate