Skip to main content
Create your own

Building a Latent Diffusion Model

Hello! Welcome back to your study of generative models.

In our last lesson, we built a solid understanding of Denoising Diffusion Probabilistic Models (DDPMs). We saw how they generate high-quality images by reversing a gradual noising process. However, we also identified their main weakness: the immense computational cost of running the U-Net denoiser on high-resolution images for hundreds of steps.

Today, we'll tackle that problem head-on. Your learning outcome is to implement a Latent Diffusion Model (LDM) like the one used in Stable Diffusion. This is the architecture that made high-resolution text-to-image generation accessible.

We will learn how LDMs achieve their remarkable efficiency by moving the entire diffusion process from the high-dimensional pixel space into a much smaller, compressed latent space. To do this, we'll explore three main components:

  1. A Perceptual Autoencoder: A powerful encoder-decoder pair that learns to compress images into a meaningful latent space without losing important details.
  2. Diffusion in Latent Space: How to adapt the DDPM process from our last lesson to operate on these compact latent representations.
  3. The LDM U-Net: The specific U-Net architecture, including the crucial addition of cross-attention layers, which is the key to conditioning the model later on.

By the end of this lesson, you will understand the complete architecture of an LDM and be ready to see how it's guided by text prompts to become Stable Diffusion.

Let's start with a high-level view.

Stable Diffusion Architecture Diagram
This diagram shows the full architecture of a Latent Diffusion Model like Stable Diffusion. It consists of three main parts: 1) The VAE (Pixel Space), which compresses and decompresses images. 2) The Denoising U-Net (Latent Space), where the diffusion process happens. 3) The Conditioning block (e.g., a text encoder), which guides generation. Today, we focus on implementing the VAE and the U-Net.

1. The Need for a Better "Compression"

The core idea of an LDM is simple: instead of running the expensive diffusion process on a 512x512 pixel image, we first compress the image into a smaller latent representation (e.g., 64x64) and run the diffusion there. This drastically reduces computational demand.

This compression is handled by an autoencoder. However, a standard autoencoder trained with a simple pixel-wise loss (like L1 or L2) tends to produce blurry, overly smooth reconstructions. It averages out high-frequency details, which are precisely what make an image look sharp and realistic. This would lead to a diffusion model that generates blurry images, defeating the purpose.

The Latent Diffusion Model paper introduced a clever solution: train the autoencoder with a combination of losses to ensure it learns a perceptually rich latent space.

Stable Diffusion from Scratch in PyTorch | Unconditional Latent Diffusion Models

To understand why LDMs are necessary and how they solve the shortcomings of simple autoencoders, let's watch a few segments from the 'Stable Diffusion from Scratch' video by ExplainingAI. This part explains the motivation and the key loss functions.

Watch the following segments: Why Latent Diffusion? (03:21 - 04:30): This sets the stage by explaining the computational bottleneck of pixel-space diffusion. The Autoencoder (04:30 - 07:31): Introduces the autoencoder's role and the problem of blurry reconstructions with simple L1/L2 loss. Perceptual Loss (07:31 - 16:46): Explains how LPIPS (Learned Perceptual Image Patch Similarity) uses features from a pre-trained network (like VGG) to enforce perceptual similarity, leading to sharper images. You can skim the code details but focus on the concept. Adversarial Loss (16:46 - 19:32): Describes adding a GAN-style discriminator to further push the autoencoder to generate realistic details that can fool the discriminator.

In summary, to get a high-quality latent space, the LDM autoencoder is trained with three objectives:

  1. Reconstruction Loss (L1/L2): Ensures the overall structure of the image is correct.
  2. Perceptual Loss (LPIPS): Enforces that the reconstructed image is perceptually similar to the original, preserving textures and details.
  3. Adversarial Loss: A GAN discriminator pushes the decoder to produce realistic, "unfakeable" high-frequency details.

This combination creates a VAE that can compress an image into a smaller latent space and reconstruct it with high fidelity.

2. The Autoencoder Architecture

With a robust training objective defined, let's examine the architecture of the VAE components. Given your background in Python and PyTorch, you'll find that the building blocks are familiar. The encoder and decoder in an LDM are typically symmetric, using a series of downsampling and upsampling blocks.

Stable Diffusion from Scratch in PyTorch | Unconditional Latent Diffusion Models

Let's continue with the ExplainingAI video to see how the autoencoder is constructed. You'll notice the components are very similar to the U-Net we saw in the previous DDPM lesson.

Watch the segment on the autoencoder architecture (19:32 - 26:55). Pay attention to how it reuses ResNet and self-attention blocks for the encoder (down-blocks) and decoder (up-blocks). Note the key differences from a U-Net: no time step information and, crucially, no skip connections between the encoder and decoder.

For a concrete code implementation, the following article provides a clear, from-scratch implementation of the VAE used in Stable Diffusion.

Implementing Stable Diffusion from Scratch using PyTorch

This article by Ebad Sayed, 'Implementing Stable Diffusion from Scratch using PyTorch', provides excellent, well-commented code for the VAE. Let's study the encoder and decoder implementations.

Review the code sections 'The Encoder' and 'The Decoder'. Observe the structure: Encoder: A sequence of Conv2d, VAE_ResidualBlock, and VAE_AttentionBlock layers that progressively downsample the image (e.g., from 512x512 to 64x64) while increasing channel depth. Decoder: A symmetric structure that uses Upsample, VAE_ResidualBlock, and VAE_AttentionBlock to reconstruct the image from the latent representation. Notice how the final layers of the encoder output 8 channels, which are then split into two 4-channel tensors for the mean and log-variance of the latent distribution.

Once this powerful autoencoder is trained, we can freeze its weights. We now have two key functions:

  • encode(image) -> latent
  • decode(latent) -> image

These allow us to move freely between the pixel and latent spaces.

3. Diffusion in Latent Space: The U-Net

Now for the main event. We take the DDPM logic from our previous lesson and apply it entirely within the latent space.

  1. Forward Process: Instead of adding noise to an image , we first encode it to get its latent representation . Then we apply the noising process to to get a noisy latent .
  2. Training: We train a U-Net, , to predict the noise from the noisy latent and the timestep . The loss function is the same MSE loss as before, just on latent variables:

This is elegantly simple. The core diffusion algorithm doesn't change, only the data it operates on.

Stable Diffusion from Scratch in PyTorch | Unconditional Latent Diffusion Models

The ExplainingAI video shows just how minimal the changes to the DDPM training loop are once you have the trained VAE.

Watch the segment starting at 39:06. Notice how the training loop is almost identical to a standard DDPM. The only additions are encoding the image to a latent before adding noise and decoding the final latent back to an image after sampling.

The Key Architectural Change: Cross-Attention

While the diffusion process is the same, the U-Net architecture used in LDMs like Stable Diffusion has one critical addition: cross-attention layers.

In the U-Net from our DDPM lesson, we likely saw self-attention, where the model relates different parts of the image to each other. Cross-attention allows the model to relate parts of the image to an external context, such as a text prompt embedding.

Even when training an unconditional LDM, the cross-attention layers are included in the architecture. They are simply fed a null or empty context. This makes the model ready for conditional training later.

Implementing Stable Diffusion from Scratch using PyTorch

Let's return to the 'Implementing Stable Diffusion from Scratch' article to see how self-attention and cross-attention are implemented and integrated into the U-Net.

First, review the 'Attention Mechanisms' section. Understand the difference in the forward methods: SelfAttention.forward(x) takes one input, x, and computes query, key, and value from it. CrossAttention.forward(x, y) takes two inputs: x (the image representation) to form the query, and y (the context) to form the key and value. Next, look at the 'The U-Net Architecture' section, specifically the UNET_AttentionBlock. Notice that it contains both a self.attention_1 (SelfAttention) and a self.attention_2 (CrossAttention) layer. This is the core block used throughout the Stable Diffusion U-Net.

Test your understanding!

In the UNET_AttentionBlock, the input x is the latent image representation and context is the conditioning information (e.g., from a text prompt). In which attention mechanism do these two inputs interact, and why is this design so powerful?

Show answer

They interact in the Cross-Attention layer (attention_2). The latent image representation x provides the queries (Q), while the context provides the keys (K) and values (V). This design is powerful because it allows every part of the image representation (Q) to "look at" every part of the conditioning context (K) and draw information from it (V), enabling fine-grained control over the generated image based on the context.

4. The Full Pipeline

We have now defined all the pieces to build a Latent Diffusion Model. The full generation process looks like this:

  1. (Optional) Text Encoding: A text prompt is converted into an embedding vector c using a text encoder like CLIP. For unconditional generation, this context is empty.
  2. Latent Denoising: Start with a random noise tensor z_T in the latent space. Iteratively denoise it from t=T to 1 using the U-Net. At each step, the U-Net takes z_t, the timestep t, and the context c (via cross-attention) to predict the noise \epsilon_\theta.
  3. Final Decoding: Once the loop finishes, you have a clean latent representation z_0. Pass this through the VAE's decoder to get the final, full-resolution image: image = decode(z_0).

This entire workflow is beautifully orchestrated in the generate function from the article we've been studying.

Implementing Stable Diffusion from Scratch using PyTorch

Let's look at 'The Pipeline' section of the article. This generate function brings together all the components we've discussed.

Read through the generate function's logic. Trace the flow of data: The context is created by passing the prompt through the clip model. latents are initialized as random noise. The main loop iterates through timesteps, calling the diffusion (U-Net) model on the latents and context. The sampler.step function updates the latents using the model's predicted noise. Finally, the decoder is called on the final latents to produce the image.

Conclusion

Excellent work! You have now dissected the architecture that powers one of the most influential generative models today. By moving diffusion to a compressed latent space, LDMs made high-resolution image synthesis practical.

Key Takeaways:

  • Efficiency via Latent Space: LDMs dramatically reduce computational cost by performing the iterative diffusion process on small latent representations, not large pixel-based images.
  • Perceptual VAE: A high-quality VAE is essential. It's trained with a combination of reconstruction, perceptual (LPIPS), and adversarial losses to ensure the latent space retains fine details.
  • Diffusion in Latent Space: The core DDPM algorithm for training and sampling remains the same but is applied to latent vectors (z) instead of images (x).
  • Cross-Attention is Key: The U-Net in an LDM is augmented with cross-attention layers. This allows the model to incorporate external conditioning information, such as text embeddings, making it a highly controllable generator.

Preview of the next lesson:
We have built and understood the complete LDM architecture, including the cross-attention layers that are waiting for a context. In the next lesson, "Apply classifier-free guidance for conditional image generation," we will finally plug something into that context input. You will learn the elegant technique of classifier-free guidance, which allows us to steer the generation process toward a text prompt by running the U-Net in two modes—conditional and unconditional—and combining the results. This is the final step in turning our LDM into a true text-to-image model like Stable Diffusion.

Can't find a good explanation? Sign up and we'll make it for you

Sign up