Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your ownLatest Computer Vision Technologies
Module 2
CNN and Vision Transformer Backbones
Module 3
Data-Efficient and Self-Supervised Vision
Module 4
Vision-Language and Open-Vocabulary Models
Module 5
Detection, Segmentation, and Visual Grounding
Module 6
Video Understanding and Persistent Tracking
Module 7
Learned Depth, Reconstruction, and Neural Rendering
Module 8
Modern Perception for UAVs and Mobile Robots
Module 9
Embodied Vision and Vision-Language-Action Systems
Module 10
Generative Image and Video Models
Module 11
Efficient Inference and Edge Deployment
Module 12
Integrated Prototypes and Research Practice