Hello! Welcome to the next lesson in our exploration of modern Text-to-Speech architectures.
In our last session, we saw how to train a complete TTS system by combining FastSpeech 2 and HiFi-GAN using an ESPnet recipe. That system represented a common two-stage approach: a text-to-mel model followed by a separate vocoder. While effective, this separation can lead to a mismatch between the two stages and requires a complex, multi-step training process (teacher-student training).
Today, we will explore a model that elegantly solves these issues. Our learning outcome is to describe the VITS architecture, explaining how it integrates a Variational Autoencoder (VAE), normalizing flows, and a GAN-based decoder for end-to-end training. VITS (Variational Inference with Adversarial Learning for End-to-End Text-to-Speech) represents a significant leap forward by combining these powerful generative modeling techniques into a single, unified framework that learns to synthesize raw audio directly from text.
This lesson directly builds upon your foundational knowledge of GANs, VAEs, and Normalizing Flows from Module 5, showing how they are masterfully orchestrated in a state-of-the-art audio synthesis model.
1. The VITS Architecture: A High-Level View
The core idea of VITS is to create a single, end-to-end differentiable model that learns a probabilistic mapping from a text sequence to a raw audio waveform. This avoids intermediate representations like mel-spectrograms acting as a "bottleneck" between separate models.
Let's start with a look at the complete architecture.

This diagram might seem complex, but we can break it down into three main conceptual pillars that you are already familiar with:
- A Variational Autoencoder (VAE) forms the backbone, learning a latent representation of speech.
- Normalizing Flows are used to create a powerful and expressive prior distribution for this latent space, conditioned on the input text.
- A Generative Adversarial Network (GAN) is used for the final decoder, ensuring the output waveform is highly realistic.
We will now dissect each of these pillars to understand their specific role within VITS.
2. The Generative Core: A Conditional VAE
At its heart, VITS is a Conditional Variational Autoencoder (CVAE). As a refresher from Module 5, a VAE learns to encode data into a latent distribution and then decode from it. A CVAE extends this by conditioning the process on some other information—in our case, the input text.
For a deeper look at the mathematical foundation, let's turn to a concise summary.
CS 224S - Spoken Language Processing Lecture Slides
These slides from a Stanford course on Spoken Language Processing provide a clear, mathematical refresher on VAEs and CVAEs, which is the perfect context for understanding VITS.
Please review the 'Appendix' slides titled 'Background: VAEs' (slide 45) and 'Background: CVAEs' (slide 46). Focus on understanding: The goal of maximizing the evidence log p(x) and the introduction of the Evidence Lower Bound (ELBO). The two terms of the ELBO: the reconstruction term and the KL divergence term. How the ELBO changes in a CVAE to include the condition c.
In the context of VITS, the CVAE components are:
- Observed Data
x: The ground-truth raw waveform. - Condition
c: The input phoneme sequence from the text. - Latent Variable
z: A latent representation that captures the high-level acoustic features of the speech, but not the fine-grained waveform details.
The VAE framework in VITS consists of several key networks:
-
Posterior Encoder
q_φ(z|x): This network takes the real audio waveformxand encodes it into the parameters of a posterior distribution (mean and variance). This encoder is only used during training. Its purpose is to provide a "target" latent representationzthat the rest of the model must learn to predict from text. -
Prior Encoder
p_θ(z|c): This network takes the input textcand encodes it into the parameters of a prior distribution. The goal of training is to make this prior distribution as close as possible to the posterior distribution. This network is used during both training and inference. -
Decoder
p_θ(x|z): This network takes a sample from the latent spacezand decodes it back into a raw waveformx. This is used in both training (for reconstruction) and inference (for synthesis).
The training is driven by the CVAE's objective function (the ELBO):
Let's break down what this means for VITS:
- Reconstruction Loss (
L_rec): The first term, , pushes the decoder to be able to reconstruct the original waveformxfrom the latentzproduced by the posterior encoder. In practice, this is often an L1 loss on the mel-spectrograms of the real and generated audio. - KL Divergence Loss (
L_kl): The second term, , is the crucial part. It forces the distribution predicted from the text (p_θ(z|c)) to match the distribution predicted from the audio (q_φ(z|x)). This is how the model learns the alignment between text and speech without any explicit supervision!
3. Making the Prior Expressive: Normalizing Flows
In a standard VAE, the prior p(z|c) is often a simple isotropic Gaussian. However, speech is incredibly complex and varied. Forcing the rich latent space of speech to fit a simple Gaussian is too restrictive and can harm synthesis quality.
VITS solves this by using Normalizing Flows to construct a much more flexible and expressive prior distribution.
CS 224S - Spoken Language Processing Lecture Slides
Let's revisit the Stanford slides for a concise explanation of how Normalizing Flows are used in VITS.
Please read the slides 'Background: Invertible Flows' (slide 47) and 'VITS: Priors' (slide 48). Focus on: The core idea: transforming a simple distribution (e.g., Gaussian) into a complex one using a sequence of invertible functions. The change of variables formula and the importance of having a tractable Jacobian determinant. How VITS uses this technique to create a more expressive prior p_θ(z|c).
As the slides explain, a normalizing flow is a chain of invertible transformations. VITS's prior encoder doesn't just output the mean and variance for a single Gaussian. Instead, it predicts parameters for a stack of affine coupling layers (similar to those in Glow/RealNVP) which transform a simple base distribution into the powerful prior p_θ(z|c). This allows the model to capture much richer variations in speech prosody and style, conditioned on the text.
4. Achieving Realism: The GAN-based Decoder
The reconstruction loss from the VAE objective (like L1 or L2) tends to produce outputs that are mathematically close to the target but sound overly smooth or muffled to the human ear. To generate crisp, high-fidelity audio, VITS incorporates adversarial training.
The decoder p_θ(x|z) is not just any network; it's a generator in a GAN setup, heavily inspired by the HiFi-GAN model we've previously discussed.
CS 224S - Spoken Language Processing Lecture Slides
The Stanford slides also neatly summarize the adversarial training component of VITS.
Please read 'VITS: Adversarial training' (slide 16) and the appendix 'Background: GANs' (slide 49). Pay attention to: How the decoder p_θ(x|z) is treated as the generator G. The introduction of a discriminator D to distinguish real and generated waveforms. The addition of an adversarial loss (L_adv) and a feature matching loss (L_fm) to the overall objective.
This adversarial setup works as a "perceptual loss". Instead of just minimizing a mathematical distance, the generator (decoder) is forced to produce waveforms that are indistinguishable from real audio to the discriminator. The feature matching loss further stabilizes training by ensuring that the intermediate features inside the discriminator are similar for both real and fake samples.
5. Tying It All Together: Alignment and The Final Objective
We have one final piece: alignment. The text encoder outputs a representation for each phoneme, while the audio waveform has thousands of samples per second. How do we align these two sequences of different lengths?
VITS uses Monotonic Alignment Search (MAS), a dynamic programming algorithm that finds the most likely alignment between the encoded text representation and the latent representation z from the audio. This happens on-the-fly during training.
This alignment also allows VITS to train a Stochastic Duration Predictor. This small network learns to predict the duration of each phoneme from the text. At inference time (when we don't have the ground-truth audio), this predictor is used to expand the phoneme embeddings to the correct length before they are passed to the normalizing flow and decoder.
The final objective function for VITS is a combination of all these ideas:
The discriminator has its own adversarial loss, L_adv(D).
Let's refer to the Stanford slides one last time to see all the components assembled.
CS 224S - Spoken Language Processing Lecture Slides
This set of slides shows how all the pieces we've discussed fit together.
Please look over slides 10 ('VITS'), 12 ('VITS' with diagram), 17 ('VITS: Alignment'), 18 ('VITS: Duration Prediction'), and 19 ('VITS: Objective'). This will help you synthesize your understanding of how the different components (VAE, Flow, GAN, Alignment, Duration) are integrated into a single model with a combined loss function.
By optimizing this single, combined loss function, VITS learns all the necessary components simultaneously: the text-to-latent mapping, the high-fidelity waveform synthesis, and the text-audio alignment.
Conclusion
In this lesson, we deconstructed the VITS architecture, revealing it to be a brilliant synthesis of several powerful generative modeling techniques.
Key Takeaways:
- Unified Framework: VITS is a true end-to-end model that trains a text-to-waveform system in a single stage, integrating a text encoder, acoustic mapper, and vocoder.
- VAE for Alignment: It uses a Conditional VAE framework where a KL divergence loss between a text-based prior and an audio-based posterior implicitly learns the alignment between text and speech.
- Normalizing Flows for Expressivity: To capture the rich complexity of speech, VITS uses a normalizing flow to model the prior distribution
p(z|c), making it far more expressive than a simple Gaussian. - GANs for Fidelity: A HiFi-GAN-style decoder and discriminator are used to ensure the final output waveform is perceptually realistic and high-fidelity.
- End-to-End Objective: The entire model is trained by optimizing a single loss function that combines reconstruction, KL divergence, adversarial, feature matching, and duration prediction losses.
Preview of the Next Lesson:
Now that we understand the core architecture of VITS, we can explore how to extend it for more advanced applications. In the next lesson, we will investigate how VITS can be conditioned on speaker embeddings to enable multi-speaker speech synthesis and even zero-shot voice cloning, a crucial step towards building highly personalized and adaptable TTS systems.