Hello! Welcome to the first lesson of our new module, Text-to-Speech Architectures.
In the previous module, we delved deep into speech analysis, learning how models like wav2vec 2.0 and HuBERT can deconstruct an audio signal into powerful, meaningful representations for tasks like speech recognition. Now, we pivot from analysis to synthesis. We will explore how to build systems that do the reverse: take text as an input and generate high-fidelity, human-like speech.
Today, we'll start by mapping out the entire landscape. Our goal is to address the learning outcome: Describe the standard TTS pipeline: text normalization, grapheme-to-phoneme conversion, acoustic model, and vocoder. We will dissect a modern neural TTS system into its core components, understanding the specific role each one plays in transforming written words into spoken audio.
By the end of this lesson, you will have a clear mental model of the end-to-end TTS process, which will serve as the foundation for the specific model architectures we'll study next.
1. The Anatomy of a Neural TTS System
At its core, Text-to-Speech (TTS) can be thought of as the inverse problem of Automatic Speech Recognition (ASR). While ASR maps a continuous audio waveform to a sequence of discrete symbols (text), TTS maps a sequence of discrete symbols back into a continuous waveform.
Historically, TTS systems were built on complex, hand-crafted rules (formant synthesis) or by stitching together pre-recorded snippets of audio (concatenative synthesis). These methods often resulted in robotic or unnatural-sounding speech. The "neural revolution" introduced a new paradigm: training deep learning models to learn the entire mapping from text to speech from data.
To see how the modern approach works, let's watch a short introduction.
Text-to-Speech & Voice Cloning Course: Neural TTS Revolution
This video from Valerio Velardo's 'The Sound of AI' series provides an excellent overview of the standard two-stage neural TTS pipeline, which has become the blueprint for many state-of-the-art systems.
Watch the clip from 00:10:32 to 00:11:45. Focus on the three main components presented: the G2P model, the acoustic model, and the neural vocoder. Note the specific input and output of each stage.
As the video outlines, the modern TTS pipeline is typically broken down into a series of distinct stages. This modular design allows us to tackle the very complex problem of speech synthesis by breaking it into more manageable sub-problems.
Let's visualize this standard pipeline:

The pipeline consists of two main parts:
- The Frontend: This part deals with processing the raw input text and converting it into a linguistic representation suitable for the neural network. This involves two key steps we'll explore: text normalization and grapheme-to-phoneme (G2P) conversion.
- The Backend (Synthesizer): This part takes the linguistic representation and generates the audio. It is itself a two-stage process consisting of:
- An Acoustic Model, which predicts an intermediate representation of the audio, typically a mel-spectrogram.
- A Vocoder, which synthesizes the final, audible waveform from that intermediate representation.
Now, let's examine each of these components in detail.
2. The Frontend: From Raw Text to Phonemes
The first challenge in any TTS system is handling the ambiguity and inconsistency of written language. The frontend's job is to clean and structure the input text into a precise phonetic sequence.
Complete Guide to Text-to-Speech (TTS) Technology
This article from Picovoice provides a clear and comprehensive explanation of the initial stages of a TTS pipeline. We'll focus on the sections describing the text processing steps.
Read the section 'TTS Pipeline', focusing on the first three stages described: '1. Text Normalization', '2. Linguistic Analysis', and '3. Phonetic Conversion (Grapheme-to-Phoneme)'. Pay attention to the examples provided for each stage.
Let's summarize the key functions of the frontend based on this reading.
Text Normalization
This is a rule-based or model-based text cleaning process. Its goal is to disambiguate and expand non-standard words into their full, spoken form. This includes:
- Numbers:
1989-> "nineteen eighty-nine" - Currency:
$25.50-> "twenty-five dollars and fifty cents" - Abbreviations:
Dr. Smith-> "Doctor Smith" - Units:
5kg-> "five kilograms"
Without this step, the model would have no idea how to pronounce symbols like $ or abbreviations like Dr..
Grapheme-to-Phoneme (G2P) Conversion
Once the text is normalized, we need to determine its pronunciation. A grapheme is a character or letter, while a phoneme is a basic unit of sound. The G2P conversion module maps grapheme sequences to phoneme sequences.
This is a non-trivial task in languages like English, which have notoriously irregular spelling. Consider these homographs (words with the same spelling but different pronunciations and meanings):
read: Can be /riːd/ (present tense) or /rɛd/ (past tense).bow: Can be /boʊ/ (a ribbon) or /baʊ/ (to bend at the waist).
The G2P module must use context to resolve this ambiguity. This is typically done using one of three methods:
- Lexicon-based: A large dictionary that maps words to their phonetic pronunciations (e.g., the CMU Pronouncing Dictionary).
- Rule-based: A set of hand-crafted rules for pronunciation.
- Model-based: A machine learning model (often a sequence-to-sequence model) trained on a large lexicon to predict phonemes from graphemes.
The output of the frontend is a clean sequence of phonemes, ready to be fed into the acoustic model. For example:
- Input text:
Hello world! - Output phonemes (ARPAbet):
HH AH0 L OW1 . W ER1 L D .
3. The Backend: From Phonemes to Waveform
With a clean sequence of phonemes, the backend's job is to generate the audio. This is where the magic of neural synthesis happens, divided into two specialized stages: the acoustic model and the vocoder.
Text-to-Speech & Voice Cloning Course: Neural TTS Revolution
Let's return to Valerio Velardo's video to understand why this part of the pipeline is split into two models and what the specific responsibility of each model is.
Watch the segment from 00:13:06 to 00:15:15. Focus on the clear distinction made between the role of the acoustic model (handling linguistic content) and the role of the vocoder (handling high-fidelity audio generation).
The Acoustic Model
The acoustic model is the bridge between the linguistic and acoustic domains.
- Input: A sequence of phonemes from the frontend.
- Output: An intermediate acoustic representation, most commonly a mel-spectrogram.
You'll recall from Module 2 that a mel-spectrogram is a time-frequency representation of audio that is perceptually relevant to human hearing. The acoustic model's task is to predict what the mel-spectrogram for the given phoneme sequence should look like.
In doing so, the acoustic model learns to control high-level aspects of speech, including:
- Prosody: The rhythm, stress, and intonation of speech (e.g., rising pitch for a question).
- Duration: How long each phoneme should be held.
- Timing: The overall pace and pausing.
Models like Tacotron 2, which we will study next, are examples of acoustic models. They are often complex sequence-to-sequence architectures, frequently involving attention mechanisms to align the input text with the output spectrogram frames.
The Vocoder
The vocoder (a portmanteau of "voice" and "encoder") is a specialist model with one job: to synthesize a high-fidelity, raw audio waveform from the intermediate representation generated by the acoustic model.
- Input: A mel-spectrogram.
- Output: A raw audio waveform (a sequence of PCM samples).
Why the separation?
The task of generating a raw waveform is incredibly challenging. Audio is sampled at a high frequency (e.g., 16,000 or 24,000 samples per second), meaning the vocoder must generate a very long sequence of values that are coherent over both short and long time scales. A mel-spectrogram is a much more compressed, lower-temporal-resolution representation.
By splitting the task, we allow each model to specialize:
- The acoustic model handles the complex mapping from language to the "content" and "style" of speech, captured in the spectrogram.
- The vocoder focuses purely on the difficult signal processing task of "inverting" the spectrogram to produce a clean, realistic waveform, effectively acting as a "neural audio synthesizer."
This modularity allows researchers to mix and match components. You can train a new acoustic model and use it with an existing, pre-trained vocoder, or vice-versa. In later lessons, we will look at vocoders like WaveNet, WaveRNN, and HiFi-GAN, which leverage architectures like causal convolutions and GANs to achieve this.
Conclusion
In this lesson, we have constructed a high-level blueprint of a modern Text-to-Speech system. We've seen that it's not a single monolithic model but a carefully designed pipeline of specialized components working in concert.
Key Takeaways:
- Standard Pipeline: A typical neural TTS system consists of a frontend (for text processing) and a backend (for audio synthesis).
- Frontend: This stage performs text normalization to clean up raw text and grapheme-to-phoneme (G2P) conversion to translate written characters into their phonetic representations.
- Backend: This stage is a two-step process:
- The Acoustic Model maps the phoneme sequence to an intermediate acoustic representation, like a mel-spectrogram, learning the prosody and timing of speech.
- The Vocoder takes the mel-spectrogram and synthesizes the final, high-fidelity audio waveform.
- Separation of Concerns: This modular design allows each component to specialize, making the overall problem more tractable and enabling independent innovation in each area.
Preview of the Next Lesson:
Now that we understand the roles of each block in the pipeline, we are ready to open up the first major component of the backend. In the next lesson, we will perform a deep dive into the Tacotron 2 architecture, one of the most influential acoustic models. We will examine its sequence-to-sequence structure, location-sensitive attention mechanism, and how it autoregressively generates a mel-spectrogram one frame at a time.