Hello! Welcome back to our course on Audio AI.
In our last lesson, we successfully generated speech using both a fine-tuned VITS model for speakers it was trained on ("seen") and a powerful XTTS model to clone voices it had never encountered ("unseen"). This is a huge step, but it raises a critical question: how good is the speech we've generated?
This lesson addresses that question directly. We'll explore the methods used in research and industry to measure the quality of synthesized speech. Our learning outcome is to evaluate synthesized speech quality using objective (e.g., PESQ) and subjective (e.g., Mean Opinion Score) metrics.
We will cover:
- Subjective evaluation, which uses human listeners and is considered the "gold standard."
- Objective evaluation, which uses algorithms to produce a numerical score, focusing on key metrics like PESQ and STOI.
- Advanced model-based metrics, which represent the current state-of-the-art in automatic quality assessment.
Understanding these evaluation methods is fundamental for any audio AI researcher or developer, as they provide the basis for comparing models, tracking progress, and diagnosing problems.
1. Subjective Evaluation: The Human Gold Standard
Ultimately, synthesized speech is for human ears. Therefore, the most definitive way to assess its quality is to ask people what they think. This is known as subjective evaluation.
1.1 Mean Opinion Score (MOS)
The most common subjective metric is the Mean Opinion Score (MOS). In a typical MOS test, listeners rate a series of audio samples on a scale, usually from 1 to 5. The final score for a system is the average of all ratings it received.

While simple in concept, conducting a reliable MOS test is complex. To understand the nuances, pros, and cons of different subjective testing methods, please review the first part of the following presentation.
Automatic Quality Assessment for Speech and Beyond
This presentation slide deck, from a tutorial at INTERSPEECH, provides an excellent overview of both subjective and objective speech quality assessment. It's a highly relevant resource for this topic.
Please review the slides under the heading 'Subjective Quality Assessment (Listening Tests)' (from slide 5 to 19). Pay close attention to: The descriptions of MOS, Pairwise comparison, and MUSHRA tests. The pros and cons of each method. The 'Critiques of MOS' and the discussion on biases, particularly the 'Range-equalizing bias'.
1.2 Key Takeaways on Subjective Testing
From that reading, here are the crucial points to remember:
- Gold Standard: Subjective tests are the ground truth for perceived quality.
- MOS is Relative: MOS scores are heavily influenced by the context of the test (the other samples, the instructions, the listeners). You cannot meaningfully compare MOS scores from two different studies. This is a frequent mistake in interpreting results.
- Cost and Time: These tests are expensive, time-consuming, and difficult to scale, which motivates the need for automated, objective metrics.
- Beyond MOS: For fine-grained comparisons between a few systems, pairwise tests ("Which sounds better, A or B?") or MUSHRA tests are often more statistically powerful and less prone to certain biases.
2. Objective Evaluation: Algorithmic Proxies
Objective metrics are algorithms that take an audio file (or a pair of files) and output a score intended to correlate with human perception. They are essential for rapid, reproducible, and low-cost evaluation during model development.
These metrics can be broadly divided into two categories:
- Intrusive (or Reference-Based): These metrics require a clean, "ground truth" reference audio sample to compare against the synthesized one. They measure the distortion introduced by the synthesis process.
- Non-Intrusive (or Reference-Free): These metrics only analyze the synthesized audio itself to predict its quality.
Let's focus on two of the most important intrusive metrics in speech processing.
2.1 Perceptual Evaluation of Speech Quality (PESQ)
PESQ (ITU-T P.862) is one of the most widely used objective metrics. It was originally designed for evaluating speech quality in telecommunications but has been widely adopted for TTS evaluation.
PESQ works by comparing the synthesized audio to a reference audio. It's not a simple difference; it involves a complex psychoacoustic model to simulate human hearing.

The core idea of PESQ is captured by the following:
where represents the symmetric disturbance (distortions that are roughly equal in both signals) and represents the asymmetric disturbance (distortions present in one signal but not the other). The constants are chosen to map these disturbances to a score that correlates with subjective MOS ratings.
- Score Range: PESQ scores typically range from -0.5 to 4.5, with higher scores indicating better quality.
2.2 Short-Time Objective Intelligibility (STOI)
While PESQ measures overall perceptual quality, STOI is designed specifically to predict speech intelligibility. It measures how well the important temporal details of the speech are preserved.
STOI works by comparing the short-term temporal envelopes of the clean and synthesized speech. The final score is the average correlation of these envelopes across different frequency bands and time frames.
where is the correlation coefficient between the temporal envelopes of the reference and synthesized signals for the -th frequency band and -th time frame.
- Score Range: STOI scores range from 0 to 1, with higher scores indicating better intelligibility.
For a more detailed yet concise explanation of these two metrics, please review the following resource.
An optimized fixed equalizer for speech enhancement
This paper uses PESQ and STOI as its core evaluation criteria and provides excellent, clear summaries of how they work and what they measure.
Please read section '3.1 Preliminary' and its subsections '3.1.1 Perceptual evaluation of speech quality (PESQ)' and '3.1.2 A Short-Time Objective Intelligibility measure(STOI)'. Focus on understanding the core principle of each metric.
2.3 The Quality-Intelligibility Trade-off
A critical insight for any practitioner is that these metrics can be at odds with each other. A system that sounds "clean" (high PESQ) might not be perfectly intelligible (lower STOI), and vice versa. For instance, some noise reduction algorithms might remove artifacts, boosting PESQ, but also slightly smear speech consonants, hurting STOI.
The paper you just read demonstrates this phenomenon, showing that optimizing for one metric can lead to a decrease in the other. This highlights the need to use a suite of metrics to get a complete picture of a model's performance.
Other common objective metrics you will encounter include:
- Mel-Cepstral Distance (MCD): A classic metric that measures the Euclidean distance between the mel-cepstra of the reference and synthesized audio. It's simpler than PESQ but less correlated with human perception.
- Word Error Rate (WER): An ASR model transcribes the synthesized speech. The WER is the error rate of that transcription compared to the ground-truth text. This is a powerful proxy for intelligibility.
3. Advanced Objective Metrics: Model-Based Evaluation
The limitations of traditional objective metrics (imperfect correlation with human judgment, need for a reference signal) have led to a new paradigm: model-based evaluation.
The idea is to train a deep learning model to act as a surrogate for a human listener. The model takes an audio sample as input and directly predicts its MOS score. Since you have a strong background in model development, this area should be of particular interest.
Please read the following section of the presentation deck, which traces the evolution of these models.
Automatic Quality Assessment for Speech and Beyond
This section details the progression from early machine learning methods to the current state-of-the-art, which leverages the same self-supervised learning (SSL) models we use for ASR.
Please review the slides under 'Model-based evaluation' (from slide 25 to 33) and the slides on the 'VoiceMOS Challenge' (slides 40 to 45). Focus on: The architectural shift from early ML to Neural Networks (MOSNet). The modern approach of fine-tuning SSL models like Wav2Vec2 for MOS prediction (SSL-MOS, UTMOS). The importance of the VoiceMOS Challenge in benchmarking these models for generalization to out-of-domain (OOD) data.
Key developments in this area include:
- MOSNet: A pioneering model using a CNN-BLSTM architecture to predict MOS scores from a spectrogram.
- SSL-based Models (e.g., UTMOS): The current state-of-the-art. These models take a large, pre-trained speech model (like wav2vec 2.0 or HuBERT) and fine-tune it on datasets of MOS-labeled audio. They have shown remarkable performance and better generalization than previous methods.
- Non-Intrusive: A major advantage of these models is that most are non-intrusive; they don't require a reference audio, making them far more flexible.
4. Practical Tools for Evaluation
Knowing the theory is one thing; applying it is another. A growing number of toolkits are available to help you compute these metrics without having to implement them from scratch.
Automatic Quality Assessment for Speech and Beyond
The final part of the presentation provides a list of publicly available resources for speech quality assessment.
Review the slides under 'Part IV: Resources' (slides 73-81). Pay attention to the tables listing 'Public Available Metrics Hubs' like ESPnet and Amphion, and take note of the VERSA framework, which aims to be a unified, comprehensive toolkit for these tasks.
To use these in practice, you would typically write a script that loads your reference and synthesized audio files and passes them to functions from one of these toolkits. Here's what that might look like in pseudocode:
# Pseudocode for using a hypothetical evaluation toolkit
import evaluation_toolkit as et
# --- Load audio files ---
# For intrusive metrics, you need both the ground truth and the synthesized version.
reference_audio, sr = et.load_audio("ground_truth_speaker_p225_001.wav")
synthesized_audio, sr = et.load_audio("synthesized_output_p225_001.wav")
# --- Calculate intrusive metrics ---
pesq_score = et.calculate_pesq(reference_audio, synthesized_audio, sample_rate=sr)
stoi_score = et.calculate_stoi(reference_audio, synthesized_audio, sample_rate=sr)
# --- Calculate non-intrusive, model-based metrics ---
# These only need the synthesized audio.
predicted_mos = et.predict_mos(synthesized_audio, sample_rate=sr, model="UTMOS")
# --- Print results ---
print(f"PESQ: {pesq_score:.2f}")
print(f"STOI: {stoi_score:.2f}")
print(f"Predicted MOS (UTMOS): {predicted_mos:.2f}")
Toolkits like ESPnet and VERSA are designed for exactly this kind of large-scale evaluation, allowing you to compute dozens of metrics across your entire test set with a single command.
Conclusion
Congratulations! You now have a solid understanding of how to measure the quality of synthesized speech. This is an indispensable skill for developing and improving TTS systems.
Key Takeaways:
- Subjective vs. Objective: Subjective evaluation (like MOS) is the ground truth but is costly and context-dependent. Objective metrics are algorithmic, scalable, and reproducible substitutes.
- No Single Metric: There is no single perfect metric. A good evaluation uses a suite of metrics. PESQ measures overall perceptual distortion, STOI measures intelligibility, and model-based predictors like UTMOS estimate subjective quality.
- Reference vs. Reference-Free: Intrusive metrics (PESQ, STOI) require a ground-truth reference audio, limiting their use cases. Non-intrusive model-based metrics are more flexible as they only need the synthesized audio.
- State-of-the-Art: Modern evaluation is moving towards large, deep learning models trained to predict human judgments, with a strong focus on their ability to generalize to new types of speech and distortions.
Preview of the Next Lesson:
Now that we have the tools to measure quality, we can start using them to conduct experiments. In the next lesson, we will analyze the impact of dataset size and quality on voice cloning performance. We will use the metrics learned today to systematically investigate how factors like the amount of training data or the cleanliness of reference audio affect the final output of our TTS models.