Hello! Welcome back to our module on Multimodal and Cross-Domain AI.
In our last lesson, you built an image captioning model that could generate a general description of an image. We explored how an encoder-decoder framework, evolving from CNN+RNN to full Transformers, can translate visual information into a sequence of text.
Today, we'll tackle a more focused task. Instead of just describing an image, our model will answer a specific question about it. This brings us to your learning outcome for this lesson: to implement a Visual Question Answering (VQA) model. This task demands a more sophisticated fusion of vision and language, as the model must comprehend both the image content and the user's query to produce a relevant answer.
The Core Challenge of VQA
The central problem in VQA is effectively combining information from two different modalities—vision and language.
Visual Question Answering: a Survey
To start, let's get a high-level overview of the VQA pipeline. This article provides an excellent survey of the topic and breaks it down into four fundamental steps.
Read the 'Introduction' section. Focus on the four major steps it outlines for solving VQA problems. These steps will form the conceptual backbone of our lesson.
As the article describes, any VQA system generally involves:
- Image Featurization: Using a vision model (like a CNN or a Vision Transformer) to extract a numerical representation of the image.
- Question Featurization: Using a language model (like an LSTM or a BERT-style Transformer) to create an embedding of the question.
- Joint Feature Representation: This is the critical step where the image and question features are fused together.
- Answer Generation: Using the fused representation to produce an answer.
There are two primary ways to approach step 4, which fundamentally changes the model's architecture: we can treat VQA as either a classification problem or a generative problem. We'll explore both.
1. Approach 1: VQA as Classification
The simpler approach is to frame VQA as a classification task. We can analyze a large dataset of questions and answers, identify the most frequent answers (e.g., "yes", "no", "red", "2", "cat"), and treat each one as a distinct class. The model's job is then to predict the most likely class for a given image-question pair.
A Mid-Fusion Architecture
A common architecture for this approach involves "mid-fusion," where the features are combined before the final prediction layer.
S1 E1: Approaching Visual Question Answering (VQA) - Vision Language Modelling Series.
This video explains a straightforward VQA architecture that treats the problem as classification. It's a great practical example of the concepts we've discussed.
Watch the following segments to understand the architecture and fusion technique: Framing the Task (02:10 - 03:07): See how VQA is framed as a discriminative/classification task. Feature Extractors (04:06 - 05:27): Note the use of separate, pre-trained models (ViT for vision, BERT for text) as feature backbones. Fusion Network (07:15 - 11:45): This is the most important part. Pay close attention to how the image and text embeddings, which may have different dimensions, are first projected to a common dimension and then fused using an element-wise product. The result is then passed through fully connected layers for the final classification.
The key idea is to bring the two modalities into a shared space where they can interact meaningfully before a final decision is made. The element-wise product (image_features * text_features) is a simple but effective way to create a joint representation that captures the interaction between visual and linguistic features.
Practical Implementation with Hugging Face
While you could build this from scratch, modern deep learning workflows often leverage high-level libraries. The Hugging Face transformers library offers pre-built models specifically for this task, encapsulating the complexity of fusion and classification.
Visual Question Answering (Hugging Face Docs)
Let's look at how to implement a classification-based VQA model using the Hugging Face transformers library. This guide demonstrates a complete, practical workflow for fine-tuning a ViLT (Vision-and-Language Transformer) model.
Read the guide from the section 'Fine-tuning ViLT' down to 'Inference'. You don't need to memorize the code, but focus on understanding the workflow: The Model (ViltForQuestionAnswering): Understand that this model is designed for VQA as a classification task. Data Preparation: Notice how the answers are mapped to integer IDs to create labels for a classifier. The labels are also given 'scores' to handle ambiguity where multiple answers are plausible. Preprocessing: See how ViltProcessor is used to handle both image and text inputs simultaneously. Training: Observe the use of the high-level Trainer API to manage the fine-tuning process. This is a common pattern in modern ML engineering. Inference: See how the fine-tuned model can be used to predict an answer for a new image-question pair.
This classification approach is powerful for questions with a limited set of common answers but struggles with "long-tail" or open-ended questions that weren't seen frequently during training.
2. Approach 2: VQA as a Generative Task
To handle open-ended questions, we can treat VQA as a generative task, similar to the image captioning models you built in the last lesson. The goal is to generate the answer text token by token.
This approach typically leverages a powerful Large Language Model (LLM) as the decoder. The challenge lies in feeding both the image and question information into the LLM in a way it can understand.
An LLM-based Architecture
A popular and effective architecture works as follows:
- Image Encoder (e.g., ViT): Extracts a sequence of feature embeddings from the image.
- Projector: A small network (often just a linear layer or a shallow MLP) maps the image embeddings into the same dimensional space as the LLM's word embeddings.
- Concatenation: The projected image embeddings are prepended to the tokenized question embeddings.
- LLM Decoder (e.g., Llama, GPT): The combined sequence is fed into the LLM, which then autoregressively generates the answer.
Here's a diagram that illustrates this modern approach:

Now, let's dive into a full implementation of this architecture.
Implement and Train VLMs (Vision Language Models) From Scratch - PyTorch
The following video implements a Vision Language Model (VLM) from scratch in PyTorch, using the generative architecture we just discussed. It combines a ViT, a projector, and a Llama model.
This is a detailed code walkthrough. Your CS background will be very helpful here. Focus on these parts: Conceptual Architecture (02:34 - 04:34): First, watch this segment to get a clear mental model of how the ViT output is projected, concatenated with the tokenized question, and fed to the Llama model. Model Implementation (19:05 - 38:10): Study this core segment carefully. It walks through the VisionLanguageModel class. Pay special attention to the forward pass, where you'll see: Image embeddings being extracted from the ViT (vision_encoder). The embeddings being passed through the image_projector. Text embeddings being created from the input_ids. The crucial torch.cat operation that fuses the two modalities along the sequence dimension. The logic for handling attention masks and labels during training vs. inference. Text Generation (38:10 - 41:35): Skim this part to see how the model's generate function is implemented, which leverages the underlying LLM's own generate method to produce text.
This generative approach is more flexible and powerful, capable of producing nuanced, open-ended answers. It represents the current state-of-the-art for many VQA tasks and is the foundation for models like GPT-4V and LLaVA.
Test your understanding!
Consider the two main approaches we've discussed: classification and generation.
- In the classification approach (like the mid-fusion model), where in the architecture does the fusion of image and text features happen? What is the nature of the final output?
- In the generative approach (like the ViT+Llama model), where does the fusion happen? What is the nature of the final output?
Show answer
-
In the classification approach, fusion happens after separate encoders have processed the image and text into single feature vectors. The fusion (e.g., element-wise product) creates a single joint vector, which is then fed into a classifier. The final output is a probability distribution over a fixed set of predefined answer classes.
-
In the generative approach, "fusion" happens by concatenating the sequence of projected image embeddings with the sequence of question embeddings. This creates a longer, combined sequence that is fed into the LLM decoder. The final output is a sequence of tokens, generated autoregressively, forming a free-text answer.
Conclusion
In this lesson, you have learned how to tackle the complex task of Visual Question Answering. You've gone beyond the general descriptions of image captioning to build models that can provide specific, query-based information about visual content.
Key Takeaways:
- VQA Fundamentals: VQA requires the fusion of visual and linguistic features to produce a relevant answer.
- Two Main Paradigms: You explored the two dominant approaches to VQA:
- Classification: Treats frequent answers as classes and predicts the most likely one. Architectures often use "mid-fusion" techniques like element-wise products on feature vectors.
- Generation: Uses an LLM decoder to autoregressively generate a free-text answer. Architectures typically fuse modalities by concatenating image and text embedding sequences.
- Architectural Patterns: You analyzed concrete implementations for both paradigms, from a mid-fusion classification network to a modern generative model combining a ViT and a Llama decoder.
- Practical Workflows: You've seen how these models can be built from scratch in PyTorch or by using high-level libraries like Hugging Face
transformersfor efficient fine-tuning.
Preview of the Next Lesson:
We have now covered multimodal learning with images and text in depth. In our next lesson, we will pivot to a new modality: audio. You will learn how to process audio signals into spectrograms for deep learning models. This is the first and most critical step in applying the powerful deep learning techniques you've learned to tasks like automatic speech recognition and music analysis.