Welcome to your next lesson in our module on Emerging Architectures and Research Frontiers.
In our previous lesson on SHAP, we explored how to interpret a model's decisions by asking, "Why did you make that prediction?" We learned to attribute the output to the input features, opening up the "black box." Today, we shift our perspective from understanding to stress-testing. We will ask, "How can I make you fail?" This leads us directly to our learning outcome: Generate adversarial examples to test model robustness.
An adversarial example is a carefully modified input, often indistinguishable from the original to a human, that causes a machine learning model to make a mistake. Understanding and generating these examples is crucial for building reliable and secure AI systems.
In this lesson, we will:
- Define what adversarial examples are and establish a threat model.
- Implement the Fast Gradient Sign Method (FGSM), a foundational technique for generating adversarial examples.
- Discuss more advanced attacks like Projected Gradient Descent (PGD) and adversarial patches.
- Briefly touch upon the concept of adversarial training as a defense mechanism.
1. What are Adversarial Examples and Why Do They Matter?
You've seen how neural networks can achieve superhuman performance on tasks like image classification. However, this high performance often comes with a surprising fragility. A small, carefully crafted perturbation—what we call adversarial noise—can be added to an image, causing a model to misclassify it with high confidence.
To get a clear picture of this phenomenon, let's start with a short video.
The video 'Adversarial Robustness' from the Center for AI Safety provides an excellent introduction to adversarial examples, explaining what they are and why they are a significant concern in AI safety.
Please watch the first three minutes of the video: Introduction (00:00 - 01:45): This segment defines adversarial examples and shows how small, often imperceptible, noise can fool a classifier. Motivation (01:45 - 03:17): This part discusses why adversarial robustness is a critical research area, especially for ensuring the safety of future AI systems.
As the video explains, this isn't just a theoretical curiosity. In safety-critical applications like autonomous driving or medical diagnosis, an attacker could exploit these vulnerabilities with potentially catastrophic consequences. Imagine a stop sign being misclassified as a speed limit sign due to a small, sticker-like patch.
The Threat Model
To formalize our discussion, we use a "threat model" to define the attacker's capabilities and goals.
-
Attacker's Knowledge:
- White-box attack: The attacker has full access to the model, including its architecture and weights. They can compute gradients, which is what we'll focus on today.
- Black-box attack: The attacker can only query the model (provide an input and get an output) without knowing its internal workings.
-
Attacker's Goal:
- Untargeted Misclassification: The goal is simply to make the model's prediction wrong, regardless of the new output.
- Targeted Misclassification: The goal is to make the model predict a specific incorrect class.

Today, we will focus on the white-box, untargeted misclassification scenario, which is the easiest to understand and implement.
2. The Fast Gradient Sign Method (FGSM)
One of the first and most intuitive methods for generating adversarial examples is the Fast Gradient Sign Method (FGSM), introduced by Ian Goodfellow et al. in 2014. The core idea is brilliantly simple and leverages the same mechanism used for training: gradients.
During training, we use backpropagation to calculate the gradient of the loss function with respect to the model's weights () and take a small step in the opposite direction to minimize the loss.
In FGSM, we calculate the gradient of the loss function with respect to the input image () and take a small step in the same direction to maximize the loss.
The formula for generating an adversarial example from an original input is:
Let's break this down:
- is the loss function (e.g., cross-entropy) for model , input , and true label .
- is the gradient of the loss with respect to the input image . It tells us how to change each pixel of to increase the loss the most.
- is the sign function, which returns +1 for positive values, -1 for negative values, and 0 for zero. This gives us the direction of the gradient for each pixel.
- (epsilon) is a small scalar that controls the magnitude of the perturbation. It's our "attack budget."

To see how this is implemented in code and to reinforce the concept, let's turn to a PyTorch tutorial.
Adversarial Example Generation
The official PyTorch tutorial 'Adversarial Example Generation' provides a concise explanation of FGSM and a clear Python function for its implementation.
Please read the following two sections: Fast Gradient Sign Attack: This section explains the intuition behind FGSM and presents the famous panda-to-gibbon example. FGSM Attack (under 'Implementation'): This section provides the fgsm_attack function in Python. Notice how it directly translates the formula into code.
The use of the sign() function is a key detail. It ensures that we apply a uniform perturbation of magnitude to every pixel, pushing the image just enough in the direction that maximizes the loss, all while staying within our L-infinity norm budget. This makes the attack highly efficient.
Test your understanding!
In the FGSM formula, , why do we add the perturbation term instead of subtracting it, as we do in standard gradient descent for training?
Show answer
In standard training (gradient descent), our goal is to minimize the loss. We move in the opposite direction of the gradient, hence we subtract the gradient step. In an adversarial attack, our goal is to maximize the loss to fool the model. Therefore, we perform gradient ascent on the input, moving in the same direction as the gradient, which is why we add the perturbation term.
3. Generating Adversarial Examples in Practice
Now, let's walk through a complete example of attacking a model. We'll continue with the PyTorch MNIST tutorial, which provides all the necessary code to train a model, attack it, and visualize the results.
Adversarial Example Generation
This reading will guide you through the rest of the 'Adversarial Example Generation' tutorial, showing the full workflow from setting up the model to evaluating the attack's impact.
Please follow these sections in the tutorial. You don't need to run the code yourself right now, but focus on understanding the steps and the results. Model Under Attack: Briefly review the simple CNN architecture used. Testing Function: This is the core loop. Pay attention to how it calculates the gradient (loss.backward()), generates the perturbed image (fgsm_attack), and then re-classifies it. Run Attack: See how the test function is called for different values of epsilon. Accuracy vs Epsilon (Results): Analyze the plot showing that as epsilon increases, the model's accuracy plummets. Sample Adversarial Examples (Results): Observe the generated images. Note the trade-off: higher epsilon values are more effective at fooling the model but also create more noticeable visual distortions.
The results are striking. With a small , the model's accuracy drops significantly, even though the perturbed digits look nearly identical to the originals. As grows, the model's performance degrades to random chance, but the noise becomes visually apparent. This illustrates the fundamental trade-off an attacker must manage between attack effectiveness and perceptibility.
4. Beyond FGSM: Stronger Attacks and Defenses
FGSM is a great starting point, but it's a single-step method and often considered a relatively weak attack. More powerful methods have since been developed.
Projected Gradient Descent (PGD)
Projected Gradient Descent (PGD) is essentially an iterative version of FGSM. Instead of taking one large step, PGD takes multiple smaller steps, clipping the total perturbation after each step to ensure it remains within the -ball around the original image. This multi-step process allows it to find more effective adversarial examples and is considered a much stronger baseline for evaluating robustness.
Let's return to the 'Adversarial Robustness' video to see a formal description of the PGD attack.
Watch the segment on the PGD attack (08:29 - 10:18). Focus on the key differences from FGSM: the iterative loop and the clipping operation (script p) that keeps the perturbation within the budget.
Adversarial Patches
Another fascinating type of attack is the adversarial patch. Instead of adding low-magnitude noise across the entire image, this method involves creating a small, localized, high-visibility patch that can be placed anywhere on an image to cause a misclassification, often to a specific target class. These are particularly concerning because they can be physically printed and used in the real world.
For an in-depth look at how these are created, you can explore the UvA DL Notebooks tutorial on Adversarial Patches. It shows how a patch can be optimized over many images to become a "universal" trigger for a target class.
A Note on Defenses: Adversarial Training
So, how do we protect our models? While many defenses have been proposed, the most effective and widely studied is adversarial training. The idea is to fight fire with fire:
- Generate adversarial examples from your training data (usually with a strong attack like PGD).
- Train the model to correctly classify these adversarial examples alongside the original clean ones.
This process essentially acts as a form of data augmentation, teaching the model to be invariant to the specific kinds of perturbations generated by the attack.
The 'Adversarial Robustness' video also provides a concise explanation of adversarial training.
Watch the segment on adversarial training (10:18 - 11:20). Understand the two main steps: creating adversarial examples and then training the network to classify them correctly.
It's important to know that adversarial robustness is an ongoing "cat-and-mouse" game between attackers and defenders, and it remains a very active area of research.
Conclusion
In this lesson, we've taken our first steps into the world of adversarial machine learning. You've seen how surprisingly easy it is to break sophisticated models by exploiting the very gradients that make them so powerful.
Key Takeaways:
- Neural networks are vulnerable to adversarial examples, where small, often imperceptible input perturbations lead to incorrect classifications.
- FGSM is a fast, white-box attack that generates these examples by performing gradient ascent on the input data to maximize the loss.
- The attack's strength is controlled by epsilon (), which balances attack effectiveness against visual perceptibility.
- More powerful attacks like PGD are iterative, and different attack types like adversarial patches exist.
- Adversarial training, where models are trained on adversarial examples, is the leading defense strategy, but achieving true robustness is still an open problem.
Preview of the Next Lesson:
Having explored ways to interpret and break standard models, we will now pivot to a new data structure. In our next lesson, we will implement Graph Neural Networks (GNNs) for learning on graph-structured data. This will open up a new class of problems where relationships and connections are as important as the individual data points themselves.