Skip to main content
Create your own
Lesson illustration

Introduction to Normalizing Flows

Hello! Welcome to the fifth lesson in our module on the Foundations of Generative Modeling.

In the previous lessons, we explored the landscapes of Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs). We saw how VAEs optimize a lower bound on the data likelihood, giving us a structured latent space at the cost of sample sharpness. We also saw how GANs bypass likelihood entirely to achieve incredible realism, but with challenges in training stability and latent space control.

You now understand the core trade-offs between these two major families. Today, we introduce a third, powerful approach that addresses a key limitation of both: Normalizing Flows. These models offer a way to compute the exact data log-likelihood while maintaining a direct, invertible mapping between the data and latent space. This is a crucial concept, as components of advanced audio models like VITS use flows to transform distributions.

Your learning outcome for this lesson is to: Describe the concept of Normalizing Flows for constructing complex probability distributions from simple ones.

We will unpack the elegant mathematical principle that allows these models to "flow" from a simple, easy-to-sample distribution (like a Gaussian) to a complex one that matches your data.

1. The Core Idea: Transforming Probability Distributions

Imagine you have a large pile of sand representing a simple, known probability distribution, like a standard normal (Gaussian) distribution. Now, imagine you carefully sculpt that sand into an intricate castle. The sand castle represents the complex, high-dimensional distribution of your real-world data (e.g., all possible human speech).

Normalizing Flows: Flow and Generative Directions
This image provides a great analogy for Normalizing Flows. We start with a simple distribution (loose sand) and learn a series of invertible transformations (the sculpting process) to create a complex distribution that models our data (the sandcastle). Because the process is invertible, we can also reverse it to turn the castle back into a simple pile of sand.

Normalizing Flows formalize this idea. They learn a series of invertible and differentiable transformations that morph a simple base distribution, , into a target distribution, , that matches the data.

The key to making this work lies in a fundamental rule of probability theory: the change of variables theorem.

2. The Mathematics: Change of Variables

When you transform a random variable, its probability density function also changes. If you stretch a region of space, the probability density in that region must decrease to ensure the total probability remains 1. Conversely, if you compress a region, the density must increase.

The change of variables theorem gives us the exact formula to calculate the new density.

  • For a single variable: If we have a variable with density and an invertible function , the density of the new variable , , is given by:

    The term is the scaling factor that accounts for the stretching or compressing of the space.

  • For multiple variables: This extends to vectors and . The scaling factor becomes the absolute value of the determinant of the Jacobian matrix of the transformation. The Jacobian matrix contains all the first-order partial derivatives of the function and its determinant measures how much a local volume of space expands or contracts.

    where is the Jacobian of the inverse function .

To build a solid intuition for this crucial theorem, the following video provides an excellent step-by-step derivation, starting from the 1D case and building up to the multi-dimensional formula.

Normalizing Flows Explained | Flow Matching Part-1 | Generative AI

This video from ExplainingAI, titled 'Normalizing Flows Explained', offers a clear, geometric derivation of the change of variables theorem. It's a great resource for understanding the 'why' behind the formula.

Please watch from the beginning to 14:48. The first part (until 11:45) focuses on the intuitive 1D case, and the second part extends this to the multi-dimensional case with the Jacobian. Pay close attention to how the conservation of probability mass necessitates the scaling factor.

Because we often work with log-probabilities for numerical stability, we can write the log-likelihood as:

This equation is the heart of Normalizing Flows. It tells us that if we can compute the inverse of our transformation and the determinant of its Jacobian, we can calculate the exact log-likelihood of any data point . This is a significant advantage over VAEs and GANs.

3. Stacking Transformations: The "Flow"

A single simple transformation is not powerful enough to model a complex data distribution. The main insight of Normalizing Flows is to compose or chain a sequence of simpler invertible transformations.

Let's say we have a sequence of invertible functions .
We start with a sample from our simple base distribution .

The full transformation is . Since the composition of invertible functions is itself invertible, we can still apply the change of variables rule.

Normalizing Flow Transformation Process
This diagram illustrates how a Normalizing Flow model transforms a simple base distribution (z0) into a complex target distribution (zK = x) through a sequence of invertible functions. The probability density is progressively 'warped' at each step.

Applying the log-likelihood formula recursively, we find that the log-determinant of the full transformation is simply the sum of the log-determinants of each individual transformation. This is because , so .

This gives us the final log-likelihood objective for a Normalizing Flow:

Or, using an equivalent formulation with the forward Jacobian:

This is the objective function we maximize during training. To see the derivation laid out clearly, I recommend the following resource.

Flow-based Deep Generative Models

The blog post 'Flow-based Deep Generative Models' by Lilian Weng is a classic, concise reference on this topic. This section provides a clear derivation of the log-likelihood for a chain of transformations.

Please read the sections "Change of Variable Theorem" and "What is Normalizing Flows?". Focus on how the single-transformation formula is expanded into the sum over a sequence of K transformations.

4. The Normalizing Flow Recipe

To build a practical Normalizing Flow model, the transformations must satisfy two crucial properties:

  1. They must be easily invertible. For training, we need to go from data to latent () to compute the likelihood. For generation, we need to go from latent to data ().
  2. The determinant of their Jacobian must be easy and efficient to compute. A naive calculation of a determinant for a matrix is , which is intractable for high-dimensional data like audio or images. We need transformations whose Jacobian has a special structure (e.g., triangular or diagonal), making the determinant calculation .

The main research in Normalizing Flows focuses on designing neural network layers that satisfy these two constraints while being expressive enough to model complex data.

In the next lesson, we will explore specific examples of these transformations, such as affine coupling layers used in models like RealNVP and Glow. These clever designs ensure the Jacobian matrix is triangular, making its determinant simply the product of the diagonal elements, thus satisfying the efficiency constraint.

To get a first glimpse of how this works, the following short video segment introduces the idea.

What are Normalizing Flows?

This clip from the Ari Seff video you saw mentioned earlier introduces the 'trick' used to make Jacobian determinant computation efficient.

Watch from 06:53 to 08:28. The video introduces the concept of ensuring the Jacobian is triangular and then shows how 'coupling layers' achieve this by leaving part of the input unchanged while transforming the rest. We will dive deep into this mechanism in our next lesson.

Conclusion

You have now learned the fundamental concept behind Normalizing Flows. Unlike VAEs that optimize a bound or GANs that discard likelihood, flows provide a method for exact likelihood estimation through a series of invertible transformations.

Key Takeaways:

  • Normalizing Flows construct complex distributions by applying a sequence of invertible transformations to a simple base distribution.
  • The Change of Variables theorem is the mathematical foundation, relating the probability densities of the original and transformed variables via the Jacobian determinant.
  • This framework allows for exact and tractable computation of the log-likelihood of the data, which is used as the training objective.
  • The key design challenge is creating transformations that are both expressive and computationally efficient, meaning they are easy to invert and their Jacobian determinant is fast to compute.

Preview of the Next Lesson:

We've covered the "what" and "why" of Normalizing Flows. In the next lesson, we'll dive into the "how." We will explore the ingenious designs of invertible transformations, such as coupling and autoregressive layers, that satisfy the necessary constraints and make Normalizing Flows practical and powerful generative models.

Can't find a good explanation? Sign up and we'll make it for you

Sign up