Hello! Welcome to our next lesson on Multimodal AI.
In our last session, we explored how the Whisper model masterfully converts human speech into text (ASR). We saw how it uses a Transformer encoder-decoder architecture to process log-Mel spectrograms and generate accurate transcriptions.
Today, we will tackle the inverse problem, fulfilling the learning outcome: Build a text-to-speech (TTS) system using a deep learning approach. We'll journey from text input to a synthesized audio waveform, deconstructing the elegant architectures that make machines speak. You'll understand the components of a modern TTS pipeline, dissect a key model architecture, and learn how to both use pre-trained systems and train your own.
1. The Modern Deep Learning TTS Pipeline
Just as ASR isn't a single step, modern TTS is a multi-stage process. While ASR goes Waveform -> Spectrogram -> Text, TTS essentially reverses this flow. Let's get a high-level overview of the components that form a typical deep learning TTS system.
Let's Start The Monster Text to Speech & Voice Cloning Course
This video from 'The Sound of AI' provides an excellent conceptual map of the modern TTS landscape. It outlines the key technologies that work together to create realistic synthetic speech.
Watch the segment from 13:30 to 17:15. As you watch, focus on identifying the three main stages of a modern TTS pipeline: 1) converting text to an intermediate representation, 2) generating a spectrogram, and 3) converting that spectrogram back to a waveform. Note the names of the technologies involved, like neural vocoders and codecs.
As the video outlines, a typical deep learning TTS pipeline consists of two main models, preceded by a text-processing step:
-
Text Frontend (Text Normalization & Phonemization): Raw text is messy. It contains numbers ("1984"), abbreviations ("Dr."), and punctuation. The first step is to clean this text and convert it into a phonetic representation. For example, "The cat sat." becomes something like
[DH, AH, K, AE, T, S, AE, T, .]. This is done using a Grapheme-to-Phoneme (G2P) converter. This gives the model a consistent, unambiguous representation of the sounds to produce. -
Acoustic Model (e.g., FastSpeech 2, Tacotron 2): This is the heart of the system. It takes the sequence of phonemes and generates a log-Mel spectrogram. This is the exact same kind of spectrogram we saw Whisper use as its input. The acoustic model's job is to predict the spectral content of the speech over time, effectively translating the symbolic phonemes into a detailed acoustic plan.
-
Vocoder (e.g., HiFi-GAN, WaveNet): A spectrogram is a visual representation of sound, not sound itself. The vocoder is a separate neural network that takes the spectrogram from the acoustic model and synthesizes the final audio waveform. Its job is to "invert" the spectrogram, filling in the phase information and other details needed to create a high-fidelity, audible sound wave.
This two-stage approach (Acoustic Model + Vocoder) is a very common and successful paradigm in TTS.
2. A Deep Dive into the Acoustic Model: FastSpeech 2
The most interesting architectural innovations often happen in the acoustic model. Let's dissect FastSpeech 2, a highly influential non-autoregressive model that addresses many challenges in TTS.
Unlike autoregressive models (like the Whisper decoder) that generate output one step at a time, FastSpeech 2 generates the entire spectrogram in parallel. This makes it extremely fast for inference.
Let's look at its architecture.

This architecture can be broken down into four key parts:
-
Phoneme Encoder: A standard Transformer encoder (using feed-forward networks and self-attention, which you're familiar with) processes the input phoneme sequence. It produces a sequence of hidden states, one for each phoneme, that are contextually aware.
-
Variance Adaptor: This is the core innovation of FastSpeech 2. The phoneme sequence is a very sparse representation. It doesn't tell the model how long to hold each sound, at what pitch to say it, or with what energy. The Variance Adaptor injects this crucial prosodic information.
- Duration Predictor: It predicts the duration (in spectrogram frames) for each phoneme. This is critical for non-autoregressive models, as it solves the alignment problem: how to expand the phoneme sequence to match the length of the audio.
- Pitch Predictor: It predicts the fundamental frequency (pitch contour) for each phoneme.
- Energy Predictor: It predicts the average energy (related to loudness) of each phoneme.
-
Length Regulator: Using the predicted durations, this simple module expands the encoder's hidden states. For example, if the phoneme /k/ has a predicted duration of 4, the Length Regulator repeats the hidden state for /k/ four times.
-
Mel-Spectrogram Decoder: Another Transformer-style decoder takes the expanded and prosody-enriched sequence and generates the final mel-spectrogram.
The diagram below gives a comparative view, showing how FastSpeech 2 improves upon the original FastSpeech.

Test your understanding!
In our last lesson, the Whisper decoder used autoregressive generation. FastSpeech 2 is non-autoregressive. What component of the FastSpeech 2 architecture is the key enabler for this parallel, non-autoregressive generation, and why?
Show answer
The Duration Predictor (within the Variance Adaptor) is the key. In an autoregressive model, the decision to stop generating frames for one sound and start the next is learned implicitly. The model just keeps generating until it predicts an "end" token or a new phoneme. In a non-autoregressive model, you need to know the full length of the output in advance. The Duration Predictor explicitly provides this information, allowing the model to allocate the correct number of frames for each phoneme and generate the entire spectrogram in one go.
3. Practical Application: From Theory to Code
Now that we understand the components, let's see how to build a TTS system in practice. We'll explore two approaches: using pre-trained models for immediate results, and training a model on custom voice data.
Approach 1: Assembling a Pre-trained System
Libraries like SpeechBrain and Hugging Face Transformers provide pre-trained models that we can easily assemble. The following example uses SpeechBrain to implement the two-stage pipeline we just discussed, combining a FastSpeech2 acoustic model with a HiFiGAN vocoder.
speechbrain/tts-fastspeech2-ljspeech
This model card for a SpeechBrain FastSpeech2 model provides a clear, concise code example for performing TTS.
Read the introductory text and the Python code block under the 'Perform Text-to-Speech (TTS) with FastSpeech2' section. Notice how two separate models, fastspeech2 and hifi_gan, are loaded and then chained together: encode_text (text-to-spectrogram) followed by decode_batch (spectrogram-to-waveform).
Here's the core logic from that resource, which you can run in a Python environment:
# You may need to install SpeechBrain first: pip install speechbrain
import torchaudio
from speechbrain.inference.TTS import FastSpeech2
from speechbrain.inference.vocoders import HIFIGAN
# 1. Load the two main components: Acoustic Model and Vocoder
fastspeech2 = FastSpeech2.from_hparams(source="speechbrain/tts-fastspeech2-ljspeech")
hifi_gan = HIFIGAN.from_hparams(source="speechbrain/tts-hifigan-ljspeech")
# Input text
input_text = "Hello, this is a test of a text to speech system."
# 2. Run the Acoustic Model (text -> spectrogram)
# This also returns predicted durations, pitch, and energy
mel_output, durations, pitch, energy = fastspeech2.encode_text([input_text])
# 3. Run the Vocoder (spectrogram -> waveform)
waveforms = hifi_gan.decode_batch(mel_output)
# 4. Save the final audio
torchaudio.save('example_tts.wav', waveforms.squeeze(1), 22050)
print("TTS audio saved to example_tts.wav")
More recent models, like VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech), integrate the acoustic model and vocoder into a single end-to-end network. This simplifies the inference process, as you only need to interact with one model.
The Hugging Face documentation for VITS shows how to use this powerful end-to-end model. It combines the stages we discussed into a single network, simplifying the user experience.
Read the introduction to VITS and look at the simple code example using the pipeline function. Compare the simplicity of this code to the two-model FastSpeech2 example. This demonstrates the trend towards more integrated, end-to-end architectures.
Using the Hugging Face pipeline, generating speech with VITS is even more straightforward:
# You will need transformers, torch, and scipy: pip install transformers torch scipy
from transformers import pipeline
from scipy.io.wavfile import write
# The pipeline handles loading the model and all pre/post-processing
synthesiser = pipeline("text-to-speech", "facebook/mms-tts-eng")
speech = synthesiser("Hello, this is a test of an end-to-end speech system.")
# The pipeline output is a dictionary containing the audio and sampling rate
write("example_vits.wav", rate=speech["sampling_rate"], data=speech["audio"])
print("TTS audio saved to example_vits.wav")
Approach 2: Training a Custom Voice Model (Voice Cloning)
What if you want the TTS system to speak in a specific voice, maybe even your own? This is known as voice cloning, and it involves training a TTS model on a custom dataset. This is the ultimate way to "build a TTS system."
The process is a classic supervised machine learning workflow:
- Data Collection: Record short audio clips of the target voice.
- Data Preparation: Transcribe each audio clip precisely. You need to create a manifest file (like a CSV or text file) that maps each audio file to its exact text content.
- Preprocessing: Ensure all audio files have the same format (e.g., sample rate, bit depth, single channel/mono).
- Training: Fine-tune a pre-trained TTS model (like Tacotron 2 or FastSpeech 2) on your custom audio-text pairs.
- Inference: Use the newly trained model to synthesize speech in the target voice.
Voice Cloning Made Simple Learn to Use Tacotron2 for TTS Voice Models
This video by Rasmurtech provides a practical, no-frills walkthrough of this entire process, from recording audio to training a Tacotron 2 model on Google Colab to clone his own voice.
You don't need to watch the entire video in detail, but skim through these sections to appreciate the practical steps involved: Data Prep (01:19 - 06:08): See how audio files are recorded, named, and paired with transcripts in a specific format. Training (14:00 - 19:52): Observe the process of setting up a Colab notebook, loading the data, and running the training loop. Notice the mention of the loss value decreasing (e.g., to below 0.3). Synthesis & Evaluation (19:52 - 25:58): This is a key part. Listen to the difference in quality between a model trained on 10 clips vs. 164 clips. It's a powerful demonstration of the importance of data quantity.
The video makes a crucial point: the quality and naturalness of the synthesized voice are directly proportional to the amount and quality of the training data. A model trained on just 25 short clips will sound robotic and garbled, while one trained on hours of clean audio can be remarkably realistic.
Conclusion
In this lesson, you've completed the other half of the speech AI equation. We've gone from text to a rich acoustic plan (spectrogram) and finally to an audible waveform.
Key Takeaways:
- Deep learning TTS systems typically use a two-stage pipeline: an Acoustic Model (text-to-spectrogram) and a Vocoder (spectrogram-to-waveform).
- Models like FastSpeech 2 use a non-autoregressive Transformer architecture with a Variance Adaptor to predict and control prosodic elements like duration, pitch, and energy, enabling fast and high-quality synthesis.
- More modern architectures like VITS offer an end-to-end solution, simplifying the pipeline by integrating the acoustic model and vocoder.
- "Building" a TTS system can mean either assembling pre-trained components for inference or undergoing the full supervised learning process of collecting, preparing, and training on a custom audio-text dataset to clone a specific voice.
- For custom training, the quantity and quality of the dataset are the most critical factors determining the final output quality.
Preview of the Next Lesson:
We've now covered foundational models for understanding and generating both text and speech. In the next module, "Specialized Generative Applications," we will push these concepts further. Our first lesson will be to Fine-tune a language model for NSFW or uncensored text generation. This will take us from the public-facing, general-purpose models we've studied into the world of model specialization and the techniques used to adapt them for specific, and sometimes controversial, domains, directly addressing an interest you expressed.