VizuaraVizuara AI Pods

Modern Robot Learning

From world models to VLAs — build modern robot learning systems from scratch.

advanced~3 hours8 pods live / 14 total

Pods in this Course

ACT (Action Chunking Transformer) Policies
1

ACT (Action Chunking Transformer) Policies

Coming soon

Diffusion Policy for Visuomotor Control
2

Diffusion Policy for Visuomotor Control

Coming soon

Vision-Language-Action (VLA) Models
3

Vision-Language-Action (VLA) Models

Coming soon

Understanding World Models from Scratch
4

Understanding World Models from Scratch

## How AI agents learn to dream about their environment — and use those dreams to make better decisions

~3h4 notebooksCase study
Offline RL & Dataset-Driven Robot Learning
5

Offline RL & Dataset-Driven Robot Learning

Coming soon

Behavior Cloning at Scale (RT-X Style)
6

Behavior Cloning at Scale (RT-X Style)

Coming soon

Foundation Models for Embodied Intelligence
7

Foundation Models for Embodied Intelligence

Coming soon

Understanding JEPA from Scratch
8

Understanding JEPA from Scratch

Understanding JEPA from Scratch

~5h9 notebooksCase study
Action Chunking Transfomers
9

Action Chunking Transfomers

Action Chunking Transfomers - The first SOTA Imitation Learning Method

~3h6 notebooksCase study
Diffusion Policy for Robotics
10

Diffusion Policy for Robotics

Diffusion Policy for Robotics - The second SOTA Imitation Learning Method

~3h6 notebooksCase study
SmolVLA
11

SmolVLA

SmolVLA - The most efficient VLA

~3h5 notebooksCase study
Pi0 - Our first Vision Language Action Model
12

Pi0 - Our first Vision Language Action Model

Pi0 - Our first Vision Language Action Model

~3h5 notebooksCase study
What are Vision Language Action Models?
13

What are Vision Language Action Models?

What are Vision Language Action Models?

~3h5 notebooksCase study
World Models Inside a Robot Policy: VLAs that Dream
14

World Models Inside a Robot Policy: VLAs that Dream

Adding a future-prediction dream step inside a vision-language-action policy: DreamVLA-style future visual state prediction before action selection (arXiv 2507.04447), latent world models in VLAs (VLA-JEPA, arXiv 2602.10098), and a from-scratch build unifying a small VLA policy with a learned world model: imagine consequences, then act.

~5h9 notebooksCase study