Hello! In our last lesson, we closed the loop on the fine-tuning process by learning how to evaluate our models for quality and iterate on them. We saw that evaluation is a crucial engineering discipline, not just an afterthought.
Today, we'll explore another critical dimension of model evaluation: safety and robustness. This lesson directly addresses the learning outcome: Understand content filtering in diffusion models and the methods used to bypass them.
We will dissect the mechanisms that developers build into models like Stable Diffusion to prevent the generation of certain types of content, particularly Not-Safe-For-Work (NSFW) material. Then, drawing on your background in AI and computer science, we will delve into the technical methods—often framed as adversarial attacks—used to circumvent these filters. This is an ongoing cat-and-mouse game at the frontier of AI safety research, and understanding both sides is essential for a complete picture of the field.
1. The Censors: How Content Filtering Works in Diffusion Models
Before we can bypass a system, we must understand its architecture. Developers of large-scale Text-to-Image (T2I) models implement safety measures to prevent misuse, such as generating violent, hateful, or non-consensual explicit imagery. These defenses generally fall into three categories.
MMA-Diffusion: MultiModal Attack on Diffusion Models
To start, let's get a formal overview of the defensive measures used in T2I models. The following paper, 'MMA-Diffusion', provides an excellent summary of the security landscape.
Please read the section titled '2. Related Work', focusing on the subsection 'Defensive methods'. This will introduce the three main strategies: AI moderators (prompt filters), post-hoc safety checkers, and concept-erased diffusion.
Let's break down these three defensive lines:
a) Prompt Filters (The Gatekeepers)
This is the first line of defense. The model or service inspects the user's text prompt before generation begins.
- Simple Keyword Filters: The most basic approach is to maintain a blocklist of "sensitive words." If your prompt contains a word from this list (e.g.,
nude,bloody), the system rejects it outright. - AI Moderators: More advanced services like Midjourney use a sophisticated classifier model (an "AI moderator"). This model reads the entire prompt and determines if its intent is to generate forbidden content, even if no specific keywords are used.
b) Post-Hoc Safety Checkers (The Bouncers)
This defense operates after an image has been generated but before it's shown to the user.
- The standard Stable Diffusion release includes a safety checker. If it flags the generated image as NSFW, it replaces the output with a black image.
- How it works: These checkers are typically vision models (like CLIP's vision encoder). They convert the generated image into an embedding vector. This vector is then compared to a set of pre-computed embedding vectors representing forbidden concepts (e.g.,
pornography,violence). If the cosine similarity between the image embedding and any forbidden concept embedding exceeds a certain threshold, the image is flagged.
c) Concept Erasure (The "Brainwashing")
This is the most advanced and invasive form of filtering. Instead of bolting on a filter, developers modify the diffusion model itself to "forget" or suppress its ability to generate specific concepts.
- Fine-tuning: A model can be fine-tuned on a dataset where NSFW concepts are explicitly associated with "safe" outputs, teaching it to avoid them. Examples include methods like ESD (Erasing Stable Diffusion).
- Inference-time Guidance: During the denoising process, the model can be guided away from the embeddings of unsafe concepts. This is the principle behind Safe Latent Diffusion (SLD).
These defenses create a challenging environment for users wanting to generate content that falls into broad or subjective "NSFW" categories, which brings us to the counter-strategies.
2. The Jailbreakers: Adversarial Attacks to Bypass Filters
"Jailbreaking" a model involves crafting inputs that trick its safety mechanisms. These techniques are active areas of security research, framed as "adversarial attacks." Given your technical background, you'll recognize these as optimization problems. We'll examine two primary attack modalities: text and image.
a) Text-Modal Attacks: Tricking the Gatekeeper
These attacks aim to create an adversarial prompt p_adv that passes the prompt filter but still guides the diffusion model to generate the desired NSFW content.
The core idea is to find a prompt p_adv that looks innocent to a filter, but whose text embedding T(p_adv) is very close to the embedding of a forbidden prompt T(p_tar). Since the diffusion U-Net only sees the embedding, it generates the forbidden content.
Let's explore two sophisticated methods for achieving this.
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
First, let's study the 'Jailbreaking Prompt Attack' (JPA). This paper introduces a clever and intuitive method for creating a 'concept embedding' that can be added to any prompt to make it NSFW.
Please read the following sections: Section 4: JPA: Jailbreaking Prompt Attack: Focus on how the attack works. Pay close attention to Equation (1), which defines a concept vector r by subtracting antonym embeddings (e.g., T("nude") - T("clothed")). Understand that this vector r is then added to the target prompt's embedding (Equation 2). The goal is to find a prefix of seemingly random tokens that can recreate this new, NSFW embedding. Section 3.2: Insights: This provides the high-level intuition for the attack. Section 5.3, Controllable NSFW Concept Rendering: Note how the scalar λ in Equation (2) allows for precise control over the degree of the NSFW concept.
The JPA method is powerful because it isolates the "essence" of a concept like nudity into a single vector. By adding this vector to a normal prompt embedding, you effectively inject the NSFW concept. The hard part, which the paper solves with gradient-based optimization, is finding a sequence of actual text tokens (the adversarial prefix) that produces an embedding close to this modified target. To avoid outputting sensitive words in the prefix, they use gradient masking, where the gradients for all sensitive tokens in the vocabulary are set to a large negative value, ensuring they are never chosen.
Another, similar approach is described in the MMA-Diffusion paper.
MMA-Diffusion: MultiModal Attack on Diffusion Models
The 'MMA-Diffusion' paper proposes a similar text-modal attack. Instead of creating a concept vector, it directly optimizes an adversarial prompt to match the embedding of a target NSFW prompt.
Read section 3.3. Text-Modal Attack. The key idea is the 'Semantic similarity-driven loss' (Equation 1), which aims to maximize cos(T(p_adv), T(p_tar)). Also, notice the 'Sensitive word regularization' technique, which is functionally identical to the gradient masking used in JPA.
Both methods leverage the same fundamental vulnerability: the vast, high-dimensional nature of the text embedding space. There are many different text strings that can map to semantically similar locations in this space, allowing an attacker to find a "safe" string that is a close neighbor to a "forbidden" one.
b) Image-Modal Attacks: Blinding the Bouncer
What if the prompt filter is weak, but the post-hoc image checker is strong? In this case, we can attack the image modality directly. This is particularly relevant for image editing tasks.
The goal here is to introduce an imperceptible perturbation to an input image, such that the final generated output will be misclassified as "safe" by the post-hoc checker.
MMA-Diffusion: MultiModal Attack on Diffusion Models
The MMA-Diffusion paper provides a clear explanation of this image-modal attack.
Read section 3.4. Image-Modal Attack and look at the first part of section 4.4. Multimodal Attack Results ('Evaluation on image modal attacks'). Understand that the attack adds a tiny amount of noise (δ in Algorithm 1) to the input image x_input. This noise is optimized to minimize the cosine similarity between the final output image's embedding and the safety checker's forbidden concept embeddings. The results in Figure 7 show how this bypasses the checker.
This is a classic adversarial attack, similar to those you may have studied for image classifiers. The key is that the perturbation δ is calculated using gradients backpropagated from the safety checker's loss function. It's a targeted attack designed to fool one specific component of the T2I pipeline.
When a model uses both prompt filtering and a post-hoc checker, these two methods can be combined into a multimodal attack to bypass both layers of security simultaneously.
Test your understanding!
You are trying to generate an image of a "robot soldier in a bloody battle," but your T2I service has both a prompt filter and a post-hoc safety checker.
- The prompt filter blocks the word "bloody."
- The safety checker blocks images containing gore.
Based on the attack methods we've discussed, how would you design a two-part strategy to bypass both filters?
Show answer
-
Text-Modal Attack for the Prompt Filter: You would use a technique like JPA or MMA-Diffusion. You could define a target prompt
p_tar= "robot soldier in a bloody battle" and an adversarial promptp_adv. You would then optimizep_adv(e.g., by finding an adversarial prefix) so that it doesn't contain the word "bloody" butT(p_adv)is semantically close toT(p_tar). For example, using the JPA method, you could compute a "gore" concept vector using antonyms likeT("gory") - T("inoffensive")and add it to the embedding of "robot soldier in a battle." -
Image-Modal Attack for the Safety Checker: This attack is for image editing, so let's adapt the scenario slightly. Assume you are editing an image of a "robot soldier." The image-modal attack would add a small, invisible perturbation to this input image. The perturbation would be calculated to ensure that the final output, even if it contains generated gore, will produce an embedding that the safety checker classifies as "safe."
By combining the adversarial prompt with the perturbed input image, you could bypass both security layers.
3. The Real World: The Spectrum of Censorship
The methods above describe how to attack a censored model. However, another approach is to use a model that has little to no censorship built-in from the start. Your interest in NSFW generation makes this distinction particularly relevant.
Model censorship exists on a spectrum. Some models are heavily filtered, while others are released with minimal restrictions.
RIP Stable Diffusion! BEST FREE UNCENSORED AI Model Is HERE!
For a current example, let's look at the recently released Flux model. This video from Aitrepreneur discusses its capabilities and, importantly, its level of censorship compared to competitors like Stable Diffusion 3.
Watch the segment from 14:13 to 14:54. The creator notes that Flux is 'not that censored' and can generate content that other models would block, but still has limits and cannot produce 'hardcore' NSFW images. This illustrates that 'uncensored' is often a relative term.
Creating a truly uncensored model often involves a deliberate training choice: removing the safety alignment. This is analogous to how uncensored Large Language Models (LLMs) are created.
Fully Uncensored GPT Is Here 🚨 Use With EXTREME Caution
This video about the Wizard Vicuna LLM provides a perfect analogy. While it discusses a text model, the principle is identical for image models.
Watch the first minute (00:00 - 00:56). Note the key phrase: 'responses that contained alignment or moralizing were removed.' The goal was to create a base model without built-in alignment so that it could be added later if desired. This is the core principle behind creating uncensored base models.
For diffusion models, this would mean fine-tuning a base model on a dataset that either lacks safety-aligned captions or, more directly, contains the very NSFW concepts you want to generate. This process effectively overwrites or dilutes the original safety training, and it directly connects back to the LoRA fine-tuning skills we've developed in previous lessons.
Conclusion
In this lesson, we've dissected the mechanisms of content control in modern diffusion models and the sophisticated techniques used to circumvent them. This is a crucial area of AI security, highlighting the tension between preventing misuse and enabling creative freedom.
Key Takeaways:
- Content filtering is multi-layered: Defenses include pre-generation prompt filters, post-generation safety checkers, and model-level concept erasure.
- Bypasses are adversarial attacks: These attacks exploit the high-dimensional embedding space to find inputs (text or images) that fool the safety mechanisms.
- Text-modal attacks create adversarial prompts whose embeddings mimic forbidden prompts without using sensitive keywords.
- Image-modal attacks add imperceptible noise to input images to make the final output evade post-hoc safety checkers.
- Uncensored models are often created by intentionally removing or overwriting the safety alignment data during training or fine-tuning.
Preview of the next lesson:
We are now at the end of our deep dive into generative models for computer vision. We will be shifting gears significantly as we begin the next module, Sequence Modeling with RNNs and Attention. We will go back to fundamental principles to understand how neural networks can process sequential data like text or time series. Our first lesson will introduce the basic Recurrent Neural Network (RNN) and the challenge of training it with Backpropagation Through Time (BPTT).