Skip to main content
Create your own

DDPM Forward and Reverse Processes

Hello! Welcome to your next lesson in our journey through generative models.

In our last session, we dove into advanced GAN architectures like Conditional GANs and StyleGAN, learning how to gain explicit control over generation and push image quality to photorealistic levels. We saw how architectural innovations like the U-Net generator, PatchGAN, and StyleGAN's mapping network created a new standard for generative modeling.

Today, we pivot to an entirely different and incredibly powerful paradigm: Denoising Diffusion Probabilistic Models (DDPMs). While GANs learn to generate in a single step, diffusion models learn to generate images by gradually refining them from pure noise. This iterative approach has proven to be remarkably stable and capable of producing state-of-the-art results, forming the backbone of modern marvels like DALL-E 2 and Stable Diffusion.

Your learning outcome for this lesson is to implement the forward and reverse processes of Denoising Diffusion Probabilistic Models (DDPM).

We will break this down into three core parts:

  1. The Forward Process: A fixed procedure of systematically adding noise to an image.
  2. The Reverse Process: The learned process of iteratively removing that noise.
  3. The Training Objective: A surprisingly simple objective where a neural network learns to predict the noise itself.

By the end of this lesson, you will understand the theory behind DDPMs and have a clear map for implementing one from scratch.

1. The Core Idea of Diffusion

Diffusion models consist of two opposing processes:

  • Forward Process (Diffusion): We start with a clean image from our dataset, . Over a large number of timesteps (e.g., 1000), we progressively add a small amount of Gaussian noise at each step. By the final step , the image is indistinguishable from pure Gaussian noise. This is a fixed, non-learned process.
  • Reverse Process (Denoising): We train a neural network to reverse this process. Starting with a sample of pure noise, , the model learns to iteratively denoise it, step by step, until it produces a clean image .

This elegant concept of destruction and creation is the heart of diffusion models. The video below provides a fantastic animated overview of this core idea.

Diffusion Models: DDPM | Generative AI Animated

To build a strong intuition, let's watch this short segment from Deepia's 'Diffusion Models: DDPM' video. It provides an excellent animated visualization of the forward (noising) and reverse (denoising) processes.

Watch the segment 'The core idea' (02:00 - 03:29). Focus on how the data distribution gradually transforms into a simple Gaussian distribution and how the model aims to learn the reverse of these steps.

Forward and Reverse Diffusion
A conceptual overview of the forward and reverse diffusion processes. The forward process q adds noise, while the reverse process p learns to remove it. Source: LearnOpenCV.

2. The Forward Process: A Mathematical Walkthrough

The forward process is defined as a Markov chain where we add Gaussian noise at each timestep according to a variance schedule . The distribution of a noisy image given the previous one is:

Here, is a scaling factor that keeps the variance from exploding. The values of are typically small and increase over time (e.g., from 0.0001 to 0.02).

A crucial property of this process is that we don't need to iterate times to get . We can sample directly from the original image in one step using a closed-form equation.

Let and . Then, the distribution of given is:

This equation is fundamental for efficient training. It means we can create a noisy version of an image for any timestep instantly. Using the reparameterization trick, we can write the sampling operation as:

where is a random noise sample.

The following resources provide a clear derivation of this process and show how it translates directly into code.

Diffusion Models: DDPM | Generative AI Animated

Let's watch the mathematical formulation of this forward process from the Deepia video. It explains how the variance-preserving process is constructed and leads to the final one-step sampling formula.

Watch the segment 'Forward process' (03:29 - 09:07). Pay close attention to the derivation of the formula for sampling x_t directly from x_0 and the definition of the alpha_bar term.

Denoising Diffusion Probabilistic Models (DDPM)

For a concise reference and a direct look at the PyTorch implementation, the labml.ai guide on DDPM is excellent. It shows how these mathematical constants are pre-calculated and used in the q_sample function.

Review the 'Denoise Diffusion' class implementation. Focus on: The __init__ method (lines 172-196), where the beta, alpha, and alpha_bar tensors are created. The q_xt_x0 method (lines 198-212), which calculates the mean and variance for the forward process distribution. The q_sample method (lines 214-230), which implements the reparameterization trick to get x_t from x_0 and noise epsilon.

3. The Reverse Process & The Simplified Loss

The goal of the reverse process is to learn the distribution , where represents the parameters of our neural network. The key insight from the DDPM paper is that if the steps are small enough, this reverse transition is also a Gaussian. Our network, therefore, needs to learn the mean and variance of this Gaussian.

Deriving the full training objective involves maximizing the Evidence Lower Bound (ELBO) on the log-likelihood, which can be mathematically intensive. However, it leads to a remarkably simple and intuitive result. After several steps of simplification, the authors found that instead of training the network to predict the mean of the denoised image, it's more effective to reparameterize the mean and train the network to predict the noise that was added to the image at timestep .

This transforms the training objective into a simple Mean Squared Error (MSE) loss:

Where:

  • is the true Gaussian noise added at timestep .
  • is the noise predicted by our neural network.
  • The network, , is typically a U-Net which takes the noisy image and the timestep as input.

The following video masterfully walks through this complex derivation, building intuition at each step.

Diffusion Models: DDPM | Generative AI Animated

This is the most theoretically dense part of diffusion models. Let's watch the Deepia video again as it breaks down the complex journey from the intractable log-likelihood to the simple MSE loss on the predicted noise.

Watch the sections 'Reverse Process' (09:07 - 14:20) and 'Simplifying the ELBO' (14:20 - 22:59). Don't worry about memorizing every equation. Focus on understanding the concepts: Why we need to approximate the reverse process. The role of the Evidence Lower Bound (ELBO) as a tractable objective. How the loss simplifies from a KL divergence between distributions to an MSE between their means. The final, brilliant trick of reparameterizing the mean to predict the noise epsilon instead.

Test your understanding!

Why is it so advantageous to train the model to predict the noise rather than, say, the denoised image or the mean ?

Show answer

Predicting the noise simplifies the loss function to a direct MSE between two variables that are on the same scale (both are standard Gaussian noise). This is a well-behaved and stable objective. If the model were to predict the denoised image directly, the loss would need to be re-weighted at each timestep to account for different signal-to-noise ratios, which can be more complex to tune. By predicting , the model's task is consistent across all timesteps: find the noise.

4. Implementation: Putting It All Together

Now that we have the theory for both processes, let's see how they are implemented in a full training and sampling loop. Your background in software engineering and PyTorch will make following this code-centric explanation straightforward.

The Training Algorithm

For each step in training:

  1. Load a batch of real images .
  2. Sample a random timestep for each image in the batch.
  3. Sample a batch of Gaussian noise .
  4. Create the noisy images using the q_sample function: .
  5. Pass and to the U-Net model to get the predicted noise .
  6. Calculate the MSE loss between the real noise and predicted noise .
  7. Update the model weights using backpropagation.

The Sampling (Inference) Algorithm

To generate a new image:

  1. Start with a pure Gaussian noise tensor .
  2. Iterate backwards from down to :
    a. Feed the current noisy image and the timestep into the U-Net to predict the noise .
    b. Use this predicted noise to calculate the mean of the reverse distribution.
    c. Sample from this distribution, usually by taking the mean and adding a small amount of noise (if ).
  3. The final result is the generated image.

The video below is an exceptional, detailed walkthrough of a real-world DDPM implementation. It connects every line of code back to the formulas in the papers.

Ultimate Guide to Diffusion Models | ML Coding Series | Denoising Diffusion Probabilistic Models

Let's watch Aleksa Gordić's 'Ultimate Guide to Diffusion Models'. This is a deep dive into the code, and it's perfect for seeing how the theory translates into a working implementation. We'll focus on the most critical parts.

This is a comprehensive walkthrough. Focus on the following key sections that implement the core logic: U-Net Model Architecture (22:47 - 28:00): Get a high-level overview of the U-Net structure used to predict the noise and, importantly, how the timestep t is embedded and fed into the model. Gaussian Diffusion Object (__init__) (38:04 - 43:49): Watch how all the mathematical coefficients (\alpha, eta, ar{\alpha}, etc.) from the paper are pre-calculated and stored. This is the setup for both the forward and reverse processes. Forward Process (q_sample) (52:05 - 55:45): See the direct implementation of the formula x_t = \sqrt{ar{\alpha}_t} x_0 + \sqrt{1 - ar{\alpha}_t} \epsilon. This shows how a noisy image is created for training. Reverse/Sampling Process (p_sample_loop) (1:07:35 - 1:20:00): This is the core of generation. Observe how the model iteratively predicts the noise (\epsilon_ heta) and uses it to take a step backwards, from x_t to x_{t-1}. The explanation here covers both training loss calculation and the inference loop.

Conclusion

Fantastic work today! We've unpacked the theory and implementation of Denoising Diffusion Probabilistic Models, a cornerstone of modern generative AI. You've moved beyond the adversarial training of GANs to understand a new, iterative approach to generation.

Key Takeaways:

  • Two-Process System: Diffusion models are defined by a fixed forward (noising) process and a learned reverse (denoising) process.
  • Efficient Forward Sampling: The forward process allows for efficient, one-step sampling of a noisy image for any timestep from the original image .
  • Noise Prediction Objective: The model (a U-Net) is trained on a simple objective: to predict the noise that was added to an image, using an MSE loss.
  • Iterative Generation: During inference, new images are generated by starting with pure noise and iteratively applying the trained model to denoise the image step-by-step.

Preview of the next lesson:
While DDPMs are powerful, they have one significant drawback: the diffusion process operates directly in the high-dimensional pixel space. Generating a 512x512 image requires running a large U-Net on a 512x512x3 tensor for hundreds of steps, which is computationally immense.

In our next lesson, we will tackle this challenge by exploring Latent Diffusion Models (LDMs). You will learn how to implement the architecture that made models like Stable Diffusion possible by moving the entire diffusion process from pixel space into a much smaller, more manageable latent space.

Can't find a good explanation? Sign up and we'll make it for you

Sign up