Skip to main content
Create your own
Lesson illustration

Normalizing Flows: The Power of Invertible Transformations

Hello! Welcome back to our exploration of Generative Modeling.

In our last lesson, we established the core principle of Normalizing Flows: they model complex data distributions by learning a sequence of invertible and differentiable transformations that warp a simple base distribution (like a Gaussian) into the target distribution. The key was the change of variables theorem, which allows for exact log-likelihood computation using the Jacobian determinant of these transformations.

We ended with a crucial question: how do we design transformations that are expressive enough to be useful, yet satisfy the strict constraints of being easily invertible and having a computationally cheap Jacobian determinant? A standard neural network layer won't do.

Today, we answer that question. Your learning outcome is to explain how invertible transformations are used in Normalizing Flow models. We will dissect the clever architectural designs that make Normalizing Flows practical, focusing on the most influential technique: coupling layers. This understanding is vital, as these transformations are the functional heart of advanced flow-based audio models.

1. The Trilemma of Transformation Design

As a quick recap, any function we use in our flow must satisfy three properties:

  1. Invertibility: We must be able to compute efficiently to calculate likelihood during training.
  2. Tractable Jacobian: Calculating must be fast. A naive complexity for a D-dimensional vector is prohibitive. We need an solution.
  3. Expressiveness: The transformation must be powerful enough to learn complex distributions. Ideally, it should be parameterized by a flexible function approximator like a neural network.

The challenge is that these requirements are often in conflict. A very expressive neural network is generally not invertible and has an intractable Jacobian. The solution lies not in restricting the network itself, but in cleverly designing how it's integrated into the transformation.

2. The Coupling Layer: A Breakthrough in Flow Design

The most common and impactful solution to this trilemma is the coupling layer, introduced in the NICE and RealNVP papers. The core idea is brilliantly simple:

  1. Split the input vector into two disjoint parts, and .
  2. Leave one part () unchanged. This is an identity transformation.
  3. Transform the other part () using a simple, element-wise, and easily invertible function (e.g., scaling and shifting). The crucial trick is that the parameters for this simple transformation are the output of a complex, powerful neural network that takes the unchanged part () as its input.

Let's watch a detailed walkthrough of this mechanism.

Normalizing Flows Explained | Flow Matching Part-1 | Generative AI

The following segment from the 'Normalizing Flows Explained' video by ExplainingAI clearly demonstrates the affine coupling layer, the most common type of coupling layer. It covers the forward pass, the inverse, and crucially, why the Jacobian determinant is so easy to compute.

Please watch from 21:10 to 26:58. Pay close attention to: How the input is split and how one part is used to generate the scale (s) and shift (t) parameters for the other part. The derivation of the inverse transformation. Notice that you don't need to invert the complex neural network. The structure of the Jacobian matrix and why it becomes lower triangular, leading to a trivial determinant calculation. The need for alternating masks to ensure all variables are transformed over a sequence of layers.

3. A Deeper Dive into the Mathematics

As you saw in the video, the coupling layer elegantly satisfies all our design constraints. Let's formalize this.

We'll denote the input as and the output as . We split both into parts A and B.

Forward Transformation ():
The first part is an identity mapping:

The second part undergoes an affine transformation (scale and shift), where the scale and shift are functions computed by a neural network NN:

Here, denotes element-wise multiplication. The exponential function ensures the scaling factor is always positive.

Inverse Transformation ():
Inverting is straightforward. First, we recover :

Since we now have , we can re-compute the scale and shift parameters using the same neural network:

And then simply invert the affine transformation for the second part:

Notice that we only need to evaluate NN in the forward direction. The expressiveness of NN is not constrained by invertibility.

Tractable Jacobian Determinant:
This is where the design truly shines. The Jacobian of the transformation has a special block structure:

Let's analyze each block:

  • (the identity matrix), because .
  • (a zero matrix), because does not depend on .
  • is a dense, complex matrix representing the derivatives of the affine parameters with respect to .
  • , a diagonal matrix, because the affine transformation is element-wise.

So, our Jacobian is lower triangular:

The determinant of a triangular matrix is simply the product of its diagonal elements.

The log-determinant, which we need for training, is therefore incredibly cheap to compute:

This is an operation—a simple sum—which is exactly what we need.

To solidify your understanding, I recommend reading the following section, which presents the same ideas in a clear, written format with helpful diagrams.

Tutorial 11: Normalizing Flows for image modeling

The UvA Deep Learning Course tutorial on Normalizing Flows provides an excellent textual explanation of coupling layers, reinforcing the concepts from the video.

Please read the section titled "Coupling layers". Focus on the mathematical formulation, the computation graph diagram, and how the implementation maps to the theory we've just discussed.

4. Other Examples of Invertible Transformations

While coupling layers are the most prominent example, they are not the only type of invertible transformation used in flows. The core principle of designing for invertibility and a tractable Jacobian applies to others as well.

Dequantization

Image and audio data are often stored as discrete integer values (e.g., 0-255). The change of variables formula is defined for continuous distributions. To handle this, a common first step in a flow is dequantization: transforming discrete data into a continuous space. This is itself an invertible transformation. A simple method is to add uniform noise:

This maps an integer value x to a continuous interval [x, x+1). More sophisticated invertible functions, like an inverse sigmoid, are then used to map this range to to match the support of a Gaussian base distribution. The Jacobian of these element-wise transformations is diagonal, making its determinant easy to compute. You can review this in the "Dequantization" section of the UVA tutorial (LINK) we just looked at.

Squeeze and Split Operations

In multi-scale architectures for high-dimensional data, like images or spectrograms, two other key invertible operations are used:

  • Squeeze: This operation reshapes the tensor, reducing spatial dimensions while increasing the number of channels. For example, a tensor can be reshaped into a tensor. This is a simple permutation of elements and is perfectly invertible.
  • Split: After squeezing, the channels are often split in two. One half is passed through more flow layers, while the other half is factored out and its likelihood is immediately calculated against the base distribution. This is also easily invertible (concatenation).

These operations allow the model to learn features at different scales, making it more efficient and powerful. You can find a concrete implementation in the "Squeeze and Split" section of the UVA tutorial (LINK).

Conclusion

You have now seen how Normalizing Flows are constructed in practice. The key is not to limit the power of the neural networks themselves, but to architect the transformations in a way that guarantees invertibility and a tractable Jacobian.

Key Takeaways:

  • Coupling layers are the workhorse of many Normalizing Flow models. They achieve expressiveness and efficiency by splitting the input, leaving one part unchanged while using it to parameterize a simple, invertible transformation on the other part.
  • This design results in a triangular Jacobian matrix, whose determinant can be computed in linear time () by summing the log-scaling factors.
  • By stacking and alternating coupling layers, the model can learn highly complex and non-linear transformations where all input dimensions influence each other.
  • Other operations like dequantization, squeezing, and splitting are also designed as invertible transformations that contribute to building powerful, multi-scale flow architectures.

Preview of the Next Lesson:

We have now explored the core principles of the three main families of generative models: VAEs, GANs, and Normalizing Flows. In our next and final lesson of this module, we will bring everything together. We will systematically compare and contrast GANs, VAEs, and Flow-based models, analyzing their respective strengths and weaknesses in terms of sample quality, diversity, training stability, and latent space properties. This will give you a complete framework for choosing the right generative tool for a given task.

Can't find a good explanation? Sign up and we'll make it for you

Sign up