VizuaraVizuara AI Pods

Vision Encoders: How Machines Learned to See — From Convolutions to Vision Transformers

Understanding the two paradigms of visual representation learning — from local feature extraction with CNNs to global attention with Vision Transformers.

intermediate~3 hours3 notebooksConvolution OperationConvolutional Neural NetworksVision TransformersPatch EmbeddingsSelf-AttentionImage Classification

Curator of this Module

Dr. Rajat Dandekar

Dr. Rajat Dandekar

Course Instructor

Dr. Rajat Dandekar is a researcher and educator specializing in AI/ML, with a passion for making complex concepts accessible through intuitive explanations and hands-on learning.

Checking access…

Learning Path

Article
1
Notebook 1
2
Notebook 2
3
Notebook 3
Case Study
Certificate