Ring-AllReduce, Choosing Batch Size, and TensorBoard GPU Profiling
How GPUs communicate, how to pick the right batch size, and how to profile your training runs to find bottlenecks.
beginner~5 hours3 notebooksCommunication Primitives: Broadcast, Scatter, Gather, ReduceRing-AllReduce from ScratchHow to Choose the Right Batch SizeGPU Profiling with PyTorch Profiler and TensorBoardOptimizing a Training Run: From Profile to Fix
Curator of this Module
Dr. Rajat Dandekar
Course Instructor
Dr. Rajat Dandekar is a researcher and educator specializing in AI/ML, with a passion for making complex concepts accessible through intuitive explanations and hands-on learning.
Checking access…
Learning Path
Article
1
Notebook 12
Notebook 23
Notebook 3Certificate