Skip to main content
Create your own
Lesson illustration

Understanding VAEs: Architecture, Objective, and the Reparameterization Trick

Hello! Welcome to the third lesson in our module on the Foundations of Generative Modeling.

In the last lesson, we delved into the architecture of Deep Convolutional Generative Adversarial Networks (DCGANs). We saw how specific architectural guidelines—like using strided convolutions, batch normalization, and particular activation functions—made the adversarial training process more stable and effective, enabling the generation of high-quality images. The core idea remained the adversarial game, where the model learns an implicit data distribution.

Today, we shift our focus to a different family of generative models that takes a more explicitly probabilistic approach. While these models are often introduced with image data, their core principles are fundamental to advanced audio generation models you're interested in, such as the vocoder in VITS.

Your learning outcome for this lesson is to: Explain the architecture and objective function of a Variational Autoencoder (VAE), including the reparameterization trick.

We will dissect the VAE from three perspectives you appreciate: the mathematical theory, the network architecture, and its practical implementation.

1. From Autoencoder to a Generative Model

You are likely familiar with the standard autoencoder from your coursework. It's a neural network trained to reconstruct its input. It consists of:

  1. An encoder that compresses the input data into a low-dimensional latent representation .
  2. A decoder that reconstructs the original data from .
    The network is trained to minimize a reconstruction loss, such as Mean Squared Error between and .

However, a standard autoencoder is not a good generative model. Its latent space is often disorganized and non-continuous. If you randomly sample a point from this latent space and pass it to the decoder, you'll likely get nonsensical output.

To see why, let's watch a brief animation.

Variational Autoencoders | Generative AI Animated

This video from Deepia provides an excellent animated introduction, clearly showing the limitations of a standard autoencoder's latent space for generation.

Watch from the beginning to 02:38. Pay attention to why simply sampling points from the latent space of a trained autoencoder fails to produce meaningful new data.

Variational Autoencoders (VAEs) solve this problem by introducing a probabilistic twist. Instead of mapping an input to a single point in the latent space, the VAE encoder maps it to a probability distribution. Specifically, for each input , the encoder outputs the parameters—a mean and a variance —of a Gaussian distribution. A latent vector is then sampled from this distribution.

This simple change has a profound effect:

  • It forces the latent space to be continuous and smoothly organized.
  • It allows us to frame the entire model in a rigorous probabilistic way, giving us a clear objective function to optimize.

The overall architecture looks like this:

Variational Autoencoder (VAE) Detailed Diagram
This diagram illustrates the complete VAE architecture. The encoder network q(z|x) maps an input x to the parameters of a latent distribution. A latent vector z is then sampled using the reparameterization trick and passed to the decoder network P(x|z) to reconstruct the input. The training is guided by the ELBO loss, which balances reconstruction quality and latent space regularization.

2. The Probabilistic Framework and the ELBO

Since you're interested in the underlying math, let's build the VAE from first principles.

A VAE is a probabilistic generative model. We assume our data is generated from some latent variable . The generative process is:

  1. Sample a latent vector from a simple prior distribution, . Typically, this is a standard multivariate normal distribution, .
  2. Generate the data point by sampling from a conditional distribution . This distribution is parameterized by a complex function (the decoder neural network with parameters ) that maps to the parameters of the output distribution.

Our goal is to find the parameters that maximize the likelihood of our observed data, . The likelihood of a single data point is given by marginalizing over the latent variables:

This integral is intractable because it requires integrating over all possible values of , which is computationally infeasible for a high-dimensional latent space and a complex decoder.

The Evidence Lower Bound (ELBO)

To handle this, we introduce an inference model (the encoder), denoted , with parameters . This network's job is to approximate the true but intractable posterior distribution .

Now, we can decompose the log-likelihood of our data, , as follows. Since the expectation of a constant is the constant itself, we can write:

Using the definition of conditional probability, , and introducing in the numerator and denominator:

Splitting the logarithm:

This gives us our final, crucial relationship:

The first term, , is the Evidence Lower Bound (ELBO). The second term is the Kullback-Leibler (KL) divergence between our approximate posterior and the true posterior. Since KL divergence is always non-negative (), the ELBO is a lower bound on the log-likelihood of the data.

By maximizing the ELBO, we are simultaneously:

  1. Pushing up the lower bound on the data likelihood, thus training our generative model.
  2. Minimizing the KL divergence between our encoder's approximation and the true posterior, thus training our inference model.

Deconstructing the ELBO

The ELBO itself can be rewritten into a more intuitive form:

This form reveals the two core components of the VAE's objective function (when framed as a loss function to be minimized, we just negate the ELBO):

  1. Reconstruction Loss: The first term, , measures how well the decoder can reconstruct the input given a latent sample from the encoder. If the decoder is assumed to be a Gaussian distribution, this term becomes equivalent to minimizing the mean squared error between the input and the reconstructed output .

  2. KL Divergence (Regularization): The second term, , acts as a regularizer. It measures the "distance" between the distribution produced by the encoder for a given input and the prior distribution . It forces the encoder to learn distributions that are close to a standard normal, which organizes the latent space and makes it suitable for generation.

Let's watch a video that visualizes this objective and its components.

Variational Autoencoders | Generative AI Animated

The same Deepia video provides a clear, animated explanation of the VAE's objective function, breaking down the ELBO into its reconstruction and regularization terms.

Please watch from 06:15 to 10:49. This segment covers the probabilistic setup and explains the two parts of the ELBO objective.

For a deeper dive into the mathematics, the following resources provide full derivations of both the ELBO and the analytical formula for the KL divergence term when both distributions are Gaussian.

Variational autoencoders - Matthew N. Bernstein

This blog post by Matthew N. Bernstein provides a rigorous yet clear mathematical walkthrough of VAE theory. The appendix is particularly relevant for your interest in the underlying math.

I recommend you read the sections "Using variational inference to fit the model", "Reducing the variance of the stochastic gradients", and "Viewing the VAE loss function as regularized reconstruction loss". For a step-by-step proof of the KL divergence formula, study the "Appendix".

3. Making it Trainable: The Reparameterization Trick

We have a clear objective function, but there's a major roadblock. The expectation term involves a sampling step. How can we backpropagate gradients through a random sampling operation to train the encoder's parameters ?

This is where the reparameterization trick comes in.

The trick is to re-express the random variable in a way that separates the randomness from the network parameters. Instead of directly sampling from the distribution defined by the encoder, , we do the following:

  1. Sample a random noise vector from a standard normal distribution, . This sampling is external to the network and has no parameters to train.
  2. Compute as a deterministic function of , , and : where denotes element-wise multiplication.
Reparameterizing the Sampling Layer
This diagram illustrates the reparameterization trick. On the left, the gradient cannot flow through the random sampling node. On the right, by making 'z' a deterministic function of the parameters and an external noise source 'epsilon', a differentiable path is created, allowing gradients to backpropagate to the encoder.

Now, the stochasticity is isolated in , while the path from the encoder parameters (which produce and ) to is fully deterministic and differentiable. This allows us to use standard backpropagation to train the entire VAE end-to-end.

This animation and a deeper lecture segment will solidify the concept.

Variational Autoencoders | Generative AI Animated

This final clip from the Deepia animation provides an intuitive explanation of the reparameterization trick and how it enables backpropagation.

Watch from 11:15 to 13:07. Focus on how the randomness is 'externalized' from the network.

Stanford CS229: Machine Learning | Summer 2019 | Lecture 20 - Variational Autoencoder

For a more formal and in-depth explanation, this segment from a Stanford CS229 lecture explains the reparameterization trick from a mathematical standpoint, which will align well with your preference for theory.

Watch from 01:31:25 to 01:38:20. The lecturer explains why the gradient is problematic without the trick and how rewriting the expectation in terms of a parameter-free noise variable solves the issue.

4. The VAE in Practice

Let's consolidate the architecture and training process with a look at a code implementation. The forward pass of a VAE and its loss function can be implemented in PyTorch as follows:

Variational autoencoders - Matthew N. Bernstein

The blog post by Matthew N. Bernstein, which you consulted for theory, also contains a concise PyTorch implementation of a VAE. This is a perfect example of how the math translates directly to code.

Please review the PyTorch code provided in the section "Comparing the implementation of a VAE with that of an autoencoder". Pay close attention to these three functions: reparameterize(self, mu, logvar): This is the direct implementation of the trick. forward(self, x): See how the encoder, reparameterization, and decoder are chained together. loss_function(output, x, mu, logvar): Notice how the reconstruction loss (mse_loss) and the analytical KL divergence loss are combined.

The full process for a single data point is:

  1. Encode: Pass through the encoder to get and .
  2. Reparameterize: Sample and compute .
  3. Decode: Pass through the decoder to get the reconstruction .
  4. Calculate Loss: Compute the total loss as the sum of the reconstruction loss (e.g., MSE between and ) and the KL divergence term, which can be calculated analytically from and .
  5. Backpropagate: Compute gradients and update the parameters and of the decoder and encoder.

Conclusion

In this lesson, you've learned about Variational Autoencoders, a powerful class of generative models built on the principles of variational inference.

Key Takeaways:

  • Architecture: A VAE consists of a probabilistic encoder that maps an input to a distribution in the latent space (typically a Gaussian defined by and ), and a probabilistic decoder that reconstructs the input from a latent sample.
  • Objective Function: VAEs are trained by maximizing the Evidence Lower Bound (ELBO), which is equivalent to minimizing a loss function composed of two terms: a reconstruction loss that ensures fidelity and a KL divergence term that regularizes the latent space.
  • The Reparameterization Trick: This crucial technique allows gradients to be backpropagated through the stochastic sampling step by reformulating the latent variable as a deterministic function of the encoder's outputs and a parameter-free noise variable .

Preview of the Next Lesson:

We have now studied two major families of generative models: GANs and VAEs. While both learn to generate data, they do so in fundamentally different ways, which results in distinct characteristics in their learned latent spaces. In the next lesson, we will distinguish between the latent spaces learned by GANs and VAEs, comparing their structure, continuity, and how they can be manipulated. We will also briefly introduce a third generative model family, Normalizing Flows, which offer another distinct approach to modeling complex distributions.

Can't find a good explanation? Sign up and we'll make it for you

Sign up