Hello! Let's continue our exploration of generative models.
Introduction
In our last lesson, you successfully implemented a Variational Autoencoder (VAE) from scratch. We focused on the practical "how": the architecture, the crucial reparameterization trick, and the two-part loss function combining reconstruction and KL divergence. However, we took the form of this loss function as a given.
Today, we transition from the "how" to the "why." This lesson is dedicated to deriving that loss function from first principles. Your learning outcome is to derive the Evidence Lower Bound (ELBO) objective for VAEs. We will unpack the mathematics to understand:
- Why we need an alternative to maximizing the data likelihood directly.
- How the principles of variational inference allow us to define a tractable objective.
- How this objective, the ELBO, naturally decomposes into the reconstruction and regularization terms you implemented previously.
This derivation is a cornerstone of modern generative modeling and will provide a solid theoretical foundation for your understanding.
The Generative Goal and Its Intractability
The fundamental goal of any generative model is to learn the true underlying data distribution, . A good model should assign a high probability to data points similar to those in our training set and a low probability to everything else. Therefore, the most natural objective is to maximize the log-likelihood of our training data: .
In a VAE, we model the data as being generated from a latent variable . The probability of observing a single data point is obtained by considering all possible latent variables that could have generated it. This is done by marginalizing over :
Here, is the prior distribution over the latent space (e.g., a standard normal distribution), and is the likelihood of the data given a latent variable, which is defined by our decoder network.
This integral presents a major problem. The decoder network defines a highly complex, non-linear function from to . This makes the integral analytically intractable—there's no closed-form solution. Furthermore, the latent space is often high-dimensional, making numerical approximation via sampling (e.g., Monte Carlo methods) computationally infeasible due to the curse of dimensionality.
Since we cannot optimize directly, we need a different approach. This is where variational inference comes in.
Variational Inference and the ELBO Derivation
The core idea of variational inference is to approximate the true, intractable posterior distribution with a simpler, tractable distribution . In a VAE, this is our encoder, parameterized by . We want to make as close as possible to the true posterior .
To measure the "closeness" of these two distributions, we use the Kullback-Leibler (KL) divergence. Our goal is to minimize .
Let's see how this minimization goal leads us to the ELBO. The following video provides a clear, step-by-step walkthrough of the mathematical derivation.
Variational Autoencoder - Model, ELBO, loss function and maths explained easily!
This video by Umar Jamil explains the derivation of the ELBO clearly and concisely. It connects all the theoretical pieces we've discussed so far.
Watch the section from 10:25 to 13:08. Follow along as the log-likelihood of the data is expanded and rearranged to reveal the ELBO and its relationship with the KL divergence.
Let's formalize and review the steps from the video. The derivation starts with the log-likelihood of a single data point, , and cleverly manipulates it.
-
Start with . We can express this as an expectation with respect to our approximate posterior , since does not depend on :
-
Apply Bayes' Rule. We can write . Substituting this in:
-
Multiply and Divide by . This is the key trick. We introduce our approximation into the equation:
-
Split the Logarithm. Using the property :
-
Identify the Terms. We now have two important terms:
- The first term is what we will call the Evidence Lower Bound (ELBO), or .
- The second term is precisely the definition of the KL divergence between our approximation and the true posterior, .
This gives us the fundamental identity:
Since the KL divergence is always non-negative (), we have the inequality:
This is why is called a lower bound on the log-evidence (another term for log-likelihood). By maximizing the ELBO, we are pushing up this lower bound, which in turn pushes up the log-likelihood of our data. We have successfully found a tractable proxy objective!
Test your understanding!
Given the equation , consider a fixed data point . What is the relationship between maximizing the ELBO, , and minimizing the KL divergence, ?
Show answer
For a fixed data point , its true log-likelihood is a constant value. Therefore, the equation becomes . Maximizing is thus equivalent to minimizing . This beautifully connects our practical objective (maximizing the ELBO) to our theoretical goal (making our approximation as close as possible to the true posterior ).
For a detailed textual walk-through of this derivation, the following resource is excellent.
Variational Inference & Derivation of the Variational Autoencoder (VAE) Loss Function
The article 'Variational Inference & Derivation of the Variational Autoencoder (VAE) Loss Function' by Samuel Odaibo provides a clear, step-by-step mathematical derivation of the ELBO.
Read the section titled 'VAE Objective'. This section meticulously works through the same derivation we just covered, showing each algebraic manipulation clearly. This is a great way to solidify your understanding.
Unpacking the ELBO: From Theory to Loss Function
Now, let's rearrange the ELBO into the form you saw in the last lesson. This will reveal the reconstruction and regularization terms.
The ELBO is:
(Here we've added to denote the parameters of our generative model/decoder).
Using the chain rule of probability, , we get:
Splitting the logarithm again:
Rearranging the second term gives the final, most common form of the ELBO:
This is exactly the objective we want to maximize!

- The Reconstruction Term: is the expected log-likelihood of the input data given a latent sample drawn from the encoder's output. Maximizing this term forces the decoder to become good at reconstructing the original data. In practice, this corresponds to minimizing a reconstruction loss like Binary Cross-Entropy or Mean Squared Error.
- The Regularization Term: measures the divergence between the distribution produced by the encoder for a given input, , and the fixed prior over the latent space, . Maximizing this term is equivalent to minimizing the KL divergence, which pushes the encoded distributions to be close to the prior (e.g., a standard normal ). This is what prevents the model from "cheating" by encoding each point into a separate, far-flung distribution, thus ensuring the latent space remains smooth and continuous.
Since standard optimizers perform minimization, our final loss function is simply the negative of the ELBO:
This is precisely the combination of reconstruction loss and KL divergence loss you implemented in the previous lesson.
An Alternative Derivation with Jensen's Inequality
It's often insightful to see that a fundamental result can be reached from different directions. The ELBO can also be derived using Jensen's Inequality, which applies to concave functions like the logarithm.
The video below offers this alternative perspective, which can deepen your intuition about why the ELBO is a lower bound.
How Neural Networks Handle Probabilities
For an alternative perspective, this video from Artem Kirsanov provides a beautiful derivation of the ELBO using Jensen's inequality. This approach offers a different and very elegant intuition.
Watch from 22:07 to 30:20. The first part (22:07-26:31) shows how Jensen's inequality establishes the ELBO as a lower bound on the log-likelihood. The second part (26:31-30:20) unpacks the ELBO into the accuracy (reconstruction) and complexity (regularization) terms, providing a conceptual interpretation that parallels what we've just discussed.
Conclusion
In this lesson, we have bridged the gap between the practical implementation of a VAE and its theoretical underpinnings. You have seen how a seemingly arbitrary loss function arises naturally from the principles of variational inference.
Key Takeaways:
- The goal of a generative model is to maximize the data log-likelihood , but this is intractable for VAEs due to a complex integral.
- Variational Inference provides a solution by introducing a tractable approximation to the posterior, (the encoder), and optimizing it to be close to the true posterior.
- This leads to the Evidence Lower Bound (ELBO), a tractable objective function that we can maximize as a proxy for the true log-likelihood.
- The ELBO elegantly decomposes into two parts: a reconstruction term that ensures the model can regenerate the data, and a regularization term (a KL divergence) that enforces a smooth, structured latent space suitable for generation.
- Maximizing the ELBO is equivalent to minimizing the negative ELBO, which gives us the practical loss function you've already used:
Reconstruction Loss + KL Divergence Loss.
Preview of the next lesson:
Now that you have a firm grasp of the probabilistic approach to generation with VAEs, we will explore a completely different paradigm. In the next lesson, you will build a basic Generative Adversarial Network (GAN). Instead of probabilistic inference, GANs frame generation as a two-player game between a generator and a discriminator. This will open up a new set of powerful techniques and conceptual tools.