Skip to main content
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
Audio AI Researcher & Developer
ยท
Module 8
Text-to-Speech Architectures
1
Understanding the Standard TTS Pipeline
Describe the standard TTS pipeline: text normalization, grapheme-to-phoneme conversion, acoustic model, and vocoder.
2
Tacotron 2 Architecture Explained
Describe the Tacotron 2 architecture, focusing on its encoder, location-sensitive attention, and autoregressive decoder.
3
FastSpeech 2: Non-Autoregressive TTS with Variance Adaptor
Describe the FastSpeech 2 architecture, highlighting its non-autoregressive design and variance adaptor module.
4
Autoregressive vs. Non-Autoregressive TTS: A Comparative Analysis
Compare autoregressive vs. non-autoregressive TTS models in terms of synthesis quality, speed, and controllability.
5
Autoregressive Vocoders: WaveNet and WaveRNN
Describe autoregressive neural vocoders like WaveNet and WaveRNN, focusing on their use of dilated causal convolutions.
6
Understanding the HiFi-GAN Vocoder
Describe the HiFi-GAN vocoder, including its generator and multi-scale/multi-period discriminators.
7
Synthesize Audio with HiFi-GAN
Generate an audio waveform from a mel-spectrogram using a pretrained HiFi-GAN model.
8
Building a Custom TTS System with ESPnet2
Train a complete TTS system (e.g., FastSpeech 2 + HiFi-GAN) using an ESPnet2 recipe.
Previous module
Self-Supervised Speech Representation
Next module
Advanced TTS and Voice Cloning