Skip to main content
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
Audio AI Researcher & Developer
ยท
Module 9
Advanced TTS and Voice Cloning
1
VITS: End-to-End Text-to-Speech with VAE, Flows, and GANs
Describe the VITS architecture, explaining how it integrates a VAE, normalizing flows, and a GAN-based decoder for end-to-end training.
2
Speaker Embeddings for Multi-Speaker TTS Conditioning
Explain how speaker embeddings are used to condition a TTS model for multi-speaker speech synthesis.
3
Zero-Shot Voice Cloning with Speaker Encoders
Describe the methodology behind zero-shot voice cloning using a separately trained speaker encoder network.
4
Preparing a Custom Dataset for Voice Cloning
Prepare a custom dataset for voice cloning, including audio segmentation, cleaning, and transcription.
5
Fine-tune Multi-speaker TTS with Coqui TTS
Configure and launch a fine-tuning job for a multi-speaker TTS model (e.g., VITS) using Coqui TTS.
6
Speech Synthesis with Fine-Tuned Models (Zero-Shot)
Synthesize speech using the fine-tuned model for both seen and unseen speakers (zero-shot).
7
Objective and Subjective Evaluation of Synthesized Speech
Evaluate synthesized speech quality using objective (e.g., PESQ) and subjective (e.g., Mean Opinion Score) metrics.
8
Dataset Influence on Voice Cloning
Analyze the impact of dataset size and quality on voice cloning performance.
Previous module
Text-to-Speech Architectures
Next module
Audio Language Models and S2S Translation