Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI theory, architecture, models
·
Module 20
Multimodal and Cross-Domain AI
1
Vision Transformers for Image Classification
Implement Vision Transformers (ViT) for image classification
2
CLIP: Joint Vision-Language Representation Learning
Implement CLIP for joint vision-language representation learning
3
Image Captioning with Encoder-Decoder Models
Build an image captioning model using an encoder-decoder architecture
4
Implementing a VQA Model
Implement a Visual Question Answering (VQA) model
5
Spectrogram Generation for Deep Learning
Process audio signals into spectrograms for deep learning models
6
Robust ASR with Whisper
Apply the Whisper architecture for robust automatic speech recognition (ASR)
7
Deep Learning for Text-to-Speech Synthesis
Build a text-to-speech (TTS) system using a deep learning approach
Previous module
Agentic AI Systems
Next module
Specialized Generative Applications