Hello! Welcome to the fourth lesson in our module on the Foundations of Generative Modeling.
In our previous lessons, we explored two of the most influential families of generative models. We started with Generative Adversarial Networks (GANs), understanding their adversarial training process and seeing it in action with DCGANs. Then, we shifted to Variational Autoencoders (VAEs), dissecting their probabilistic foundations, the ELBO objective function, and the crucial reparameterization trick that makes them trainable.
You now have the architectural and mathematical blueprints for both. Today, we'll place them side-by-side to compare a critical component: the latent space. The way these models learn and structure this compressed representation z has profound implications for what we can do with them. This understanding is key for your goal of working with advanced audio models, as different components of models like VITS and HiFi-GAN are built on these very principles.
Your learning outcome for this lesson is to: Distinguish between the latent spaces learned by GANs and VAEs.
We will explore how their fundamentally different objectives give rise to latent spaces with distinct properties of continuity, structure, and utility.
1. The Core Difference: Inference vs. Generation
The most fundamental distinction between the latent spaces of VAEs and GANs stems directly from their architecture and training objectives.
- VAEs have an encoder. They are trained to perform inference: mapping a data point
xto a distribution in the latent space,q_φ(z|x). The entire framework is built around maximizing a lower bound on the data's log-likelihood. - GANs do not have an encoder. The generator,
G(z), learns a direct mapping from a simple noise distributionp(z)to the data space. Its objective is not to model the data likelihood but to produce samples that can fool a discriminator.
This architectural difference is the root of all other distinctions.

Let's quickly recall their objective functions, which shape their behavior:
-
VAE Loss (minimized):
The VAE objective explicitly forces the latent space to be organized. The
D_KLterm acts as a regularizer, pushing the distributionq_φ(z|x)for every inputxto be close to a simple priorp(z)(e.g., a standard normal distribution). -
GAN Loss (min-max game):
The GAN objective has no term that explicitly regularizes the structure of the latent space
z. Any organization that arises is an indirect consequence of the generator learning to mapp(z)to the data distribution in a way that fools the discriminator.
Now, let's explore the consequences of these different approaches.
2. The VAE Latent Space: Continuous and Structured
The VAE's objective function is carefully designed to produce a well-behaved latent space.
Continuity and Smoothness
The KL divergence term forces the encoded distributions of different data points to overlap and stay close to the origin. This prevents the model from "cheating" by mapping each input to a distinct, isolated region of the latent space. The result is a continuous and smooth latent space. Any point you sample from the prior distribution p(z) is likely to be decoded into a plausible, albeit potentially blurry, output.
This smoothness is what makes VAEs excellent for interpolation. If you encode two images, x1 and x2, into their latent means z1 and z2, you can walk along the straight line between z1 and z2, decode each point along the way, and see a smooth, meaningful transition from one image to the other.

To see this in action, the following animation visualizes how the VAE organizes the latent space and enables seamless blending between images.
Variational Autoencoders | Generative AI Animated
This segment from the Deepia animation we saw previously provides a powerful visualization of the VAE's trained latent space.
Watch from 13:59 to 16:30. Observe how the latent space conforms to a normal distribution and how this enables smooth interpolation between different digits.
Interpretability and Disentanglement
A fascinating side effect of this regularization is that the dimensions of the latent space often learn to capture independent, interpretable factors of variation in the data. This property is known as disentanglement. For example, in a VAE trained on faces, one latent dimension might learn to control smile intensity, another might control head rotation, and a third might control lighting direction.
This is a powerful feature for data analysis and controllable generation.
This video from Arxiv Insights provides one of the clearest demonstrations of a disentangled latent space in a VAE.
Please watch from 09:05 to 13:03. Pay close attention to how modifying a single latent variable in the 'disentangled' VAE results in a specific, interpretable change in the output image (like changing floor color or object rotation), whereas in the standard VAE, everything gets blurry.
The Trade-off: The price for this beautifully structured latent space is often sample quality. The reconstruction loss, especially when it's a simple Mean Squared Error, encourages the decoder to average out details, leading to blurry or overly smooth samples.
3. The GAN Latent Space: Optimized for Realism
The GAN's latent space is shaped by a completely different pressure: fooling the discriminator.
High-Fidelity Samples
The adversarial objective relentlessly pushes the generator to produce samples that are indistinguishable from real data. This results in outputs that are typically much sharper and more realistic than those from a VAE.
Unstructured and Entangled
Since there is no explicit regularization on the latent space, the generator is free to learn any mapping from z to x that works. The latent space is often highly non-linear, entangled, and can have "holes". The generator might learn to map the prior p(z) onto a complex, twisted manifold in the data space.
- Interpolation: While interpolation is possible, it's less reliable than in a VAE. A linear path between two latent vectors
z1andz2may cross regions that don't map to meaningful outputs. - Entanglement: A single latent dimension often controls multiple output attributes simultaneously, making controllable generation difficult without more advanced techniques.
Mode Collapse
A notorious failure mode for GANs is mode collapse. This occurs when the generator finds a few outputs that are very effective at fooling the discriminator and maps many different input z vectors to this small set of outputs. The generator fails to capture the full diversity of the training data. This is a direct risk of the adversarial dynamic, which doesn't explicitly incentivize covering the entire data distribution.
The Invertibility Problem
As GANs lack an encoder, there's no straightforward way to find the latent code z that corresponds to a given real-world image x. This task, known as GAN inversion, requires separate, complex optimization or training an additional encoder network. This makes GANs less suitable for tasks that require editing or analyzing existing data.
4. A Modern Perspective: The Best of Both Worlds?
While the classic distinction holds, modern generative modeling, particularly in the two-stage approaches that are state-of-the-art, often blurs these lines. This is a crucial insight for someone aiming for a research career.
Let's read a perspective from a leading researcher in the field.
Generative modelling in latent space
Sander Dieleman's blog post 'Generative modelling in latent space' offers a fantastic, modern take on this topic. It moves beyond the textbook definitions to explain how these models are used in practice today.
Please read the sections titled "Curating and shaping the latent space" and the "Closing thoughts". Pay special attention to two key ideas: The argument that the 'V' in modern VAEs is often 'vestigial' because the KL term is scaled down so much, making them more like KL-regularized autoencoders. The rise of models like VQGAN, which combine an autoencoder architecture (like a VAE) with an adversarial loss (like a GAN) to get sharp reconstructions within a learned latent space.
As you read, the takeaway is that practitioners have found ways to combine the strengths of both approaches. By using an autoencoder to learn a compact representation but training it with an adversarial loss alongside a perceptual loss, models like VQGAN achieve both a structured latent space and high-fidelity reconstructions. This hybrid approach forms the foundation of powerful models like Stable Diffusion.
Summary Table
Here is a table summarizing the key distinctions we've discussed:
| Feature | Variational Autoencoder (VAE) | Generative Adversarial Network (GAN) |
|---|---|---|
| Core Architecture | Encoder-Decoder | Generator-Discriminator |
| Objective | Maximize Evidence Lower Bound (ELBO) | Minimax adversarial game |
Latent Space z |
Explicitly regularized via KL divergence | Implicitly learned; no direct regularization |
| Structure | Continuous, smooth, organized | Potentially discontinuous, "holey," and entangled |
| Primary Strength | Good for interpolation, editing, analysis (interpretable z) |
Generating high-fidelity, sharp, realistic samples |
| Primary Weakness | Often produces blurry or overly smooth samples | Prone to mode collapse; unstable training |
Inference (x -> z) |
Built-in (the encoder) | Requires separate optimization/network (GAN Inversion) |
| Typical Use in Audio | Modeling representations (e.g., in VITS vocoder) | Generating high-quality waveforms (e.g., HiFi-GAN vocoder) |
Conclusion
You can now distinguish between the latent spaces learned by VAEs and GANs, not just by their definitions but by their functional properties and the trade-offs they represent.
Key Takeaways:
- VAE latent spaces are structured for inference and control. The ELBO objective, with its KL regularization term, explicitly enforces a smooth, continuous latent space, making it ideal for tasks like interpolation and disentangled attribute manipulation. The cost is often sample realism.
- GAN latent spaces are structured for realism. The adversarial loss focuses solely on output quality, leading to sharp, high-fidelity samples. The latent space itself is often unstructured and entangled, and the training can be unstable.
- The fundamental difference is the encoder. VAEs learn a mapping from data to the latent space, while GANs only learn the reverse.
- Modern models often hybridize these ideas, using VAE-like autoencoders with GAN-like adversarial losses to get the best of both worlds.
Preview of the Next Lesson:
We have explored models that work with an approximate, tractable lower bound on the likelihood (VAEs) and models that bypass likelihood estimation entirely (GANs). What if we could design a model that computes the exact log-likelihood of the data and provides a perfectly invertible mapping between the data and the latent space? This is the promise of our next topic: Normalizing Flows. These models offer another unique approach to generative modeling with a completely different set of properties and trade-offs.