Skip to main content
Create your own
Lesson illustration

ASR Evaluation: WER and CER

Hello! Welcome to the final lesson of our module on audio data augmentation and pipelines.

In our previous lesson, we explored forced alignment, a powerful technique for synchronizing audio with text to generate precise word-level timestamps. We saw how it's crucial for training Text-to-Speech models and for detailed data analysis.

Now that we can process, augment, and annotate our audio data, a fundamental question arises: How do we measure the performance of a Speech-to-Text model? Just listening to a few outputs isn't scalable or objective. This lesson directly addresses that question.

Our learning outcome is to implement and interpret standard ASR evaluation metrics, specifically Word Error Rate (WER) and Character Error Rate (CER). We'll delve into the theory behind these metrics, learn how to calculate them, implement them in Python, and understand their limitations—a critical skill for anyone serious about ASR development and research.

1. The Need for Objective Metrics

When you train a model, you need a quantifiable way to tell if it's getting better. For Automatic Speech Recognition (ASR), the most common way to do this is to compare the model's output transcript (the hypothesis) against a perfect, human-verified transcript (the reference or ground truth).

The "distance" between these two texts is quantified by error rates. Let's explore the two industry standards.

Metrics for ASR Performance: WER and CER

This article from ApX Machine Learning provides a clear and concise introduction to the two most important metrics in ASR. We will use it as our guide for this section.

Please read the introduction and the section on 'Word Error Rate (WER)'. Focus on the definitions of the three types of errors: Substitutions, Deletions, and Insertions.

2. Word Error Rate (WER)

As the article explains, Word Error Rate (WER) is the most ubiquitous metric for ASR. It's based on the Levenshtein distance, a concept you might have encountered in your computer science studies. The Levenshtein distance measures the minimum number of single-unit edits required to change one sequence into another. For WER, these units are words.

The three types of errors are:

  • Substitutions (S): A word in the reference is replaced by a different word in the hypothesis.
    • Reference: "the cat sat"
    • Hypothesis: "the hat sat" (1 substitution)
  • Deletions (D): A word in the reference is missing from the hypothesis.
    • Reference: "the cat sat"
    • Hypothesis: "the sat" (1 deletion)
  • Insertions (I): A word appears in the hypothesis that was not in the reference.
    • Reference: "the cat sat"
    • Hypothesis: "the cat just sat" (1 insertion)

The WER Formula

The formula to calculate WER is:

Where:

  • is the number of substitutions.
  • is the number of deletions.
  • is the number of insertions.
  • is the total number of words in the reference transcript.

The process of finding the minimum number of edits (S, D, and I) is an alignment problem, typically solved with a dynamic programming algorithm.

Levenshtein Distance Calculation Explained
This diagram illustrates the dynamic programming approach to calculating the Levenshtein distance. A matrix is filled to find the minimum cost (number of edits) to transform one sequence into another. The same principle applies to WER, but the units are words instead of characters.

Let's walk through the example from the article:

  • Reference: SHOW ME THE WEATHER ()
  • Hypothesis: SHOW THE WEATHER NOW

To calculate the errors, we align them:

Reference Hypothesis Error Type
SHOW SHOW Correct
ME Deletion
THE THE Correct
WEATHER WEATHER Correct
NOW Insertion

From this alignment, we can count the errors:

  • Substitutions (S) = 0
  • Deletions (D) = 1 (the word ME was deleted)
  • Insertions (I) = 1 (the word NOW was inserted)
  • Number of words in reference (N) = 4

Plugging this into the formula:

A lower WER is better, with 0% indicating a perfect transcription.

3. Character Error Rate (CER)

While WER is excellent for many languages, it's less suitable for those that don't use spaces to separate words (like Mandarin or Japanese). In these cases, or when the accuracy of individual characters is paramount (e.g., transcribing serial numbers), we use Character Error Rate (CER).

The principle and formula are identical to WER, but the calculation happens at the character level.

Metrics for ASR Performance: WER and CER

Now, let's read the next section of the same article to understand CER.

Please read the section 'Character Error Rate (CER)'. Notice how the logic directly mirrors that of WER.

The CER Formula

Where:

  • , , and are substitutions, deletions, and insertions at the character level.
  • is the total number of characters in the reference transcript.
Calculating Character Error Rate (CER)
This image provides a clear, practical example of calculating CER. It aligns the ground truth with the model's output, counts the character-level errors (S, D, I), and plugs them into the formula.

4. Practical Implementation in Python

You rarely need to implement the Levenshtein algorithm from scratch. Several libraries do it for you. A popular, robust choice is jiwer. The Hugging Face evaluate library also provides an easy-to-use wrapper for WER and CER calculation, which is often used in training scripts.

In fact, you've already seen this in action. Let's quickly review a clip from the Whisper fine-tuning video.

Fine tuning Whisper for Speech Transcription

This clip shows how a WER metric is loaded and prepared for use in a Hugging Face trainer. It demonstrates how these metrics are integrated into a typical MLOps workflow.

Watch the section from 00:37:45 to 00:38:25. Notice how a 'word error metric' is loaded as a standard part of the training setup.

The video shows how straightforward it is to plug these metrics into a training pipeline. Now, let's see how to use a library like jiwer to perform the calculation ourselves.

Metrics for ASR Performance: WER and CER

This final section of the ApX article provides a ready-to-use Python code snippet for calculating WER and CER.

Read the section 'Calculating WER and CER in Python' and study the provided code. This is a practical skill you will use frequently.

Here's the essential code for calculating WER using jiwer.

First, install the library:

pip install jiwer

Then, you can use it in your Python script:

import jiwer




# The ground truth and the model's prediction
reference = "show me the weather"
hypothesis = "show the weather now"




# jiwer automatically handles normalization (lowercase, punctuation removal)
transformation = jiwer.Compose([
    jiwer.ToLowerCase(),
    jiwer.RemovePunctuation(),
    jiwer.RemoveMultipleSpaces(),
    jiwer.Strip(),
    jiwer.SentencesToListOfWords()
])




# Calculate the measures
output = jiwer.process(reference, hypothesis, transformation, transformation)

print(f"Word Error Rate: {output.wer:.2%}")
print(f"Substitutions: {output.s}")
print(f"Deletions: {output.d}")
print(f"Insertions: {output.i}")
print(f"Reference words (N): {output.n}")




# To calculate CER, you simply work with characters
reference_chars = list("hello")
hypothesis_chars = list("hallow")




# jiwer.process can take lists directly
output_cer = jiwer.process(reference_chars, hypothesis_chars)

print(f"\nCharacter Error Rate: {output_cer.wer:.2%}")

Running this script will produce the results we calculated manually earlier, confirming our understanding.

5. Interpreting and Using Error Rates

A WER or CER score is just a number. To make it useful, you need to understand its context and limitations.

How to Improve AI Models With Word Error Rate Metric

This article from Galileo provides excellent insights into the limitations of WER and why it should not be used in isolation.

Please read 'Step #3: Interpret and analyze the result' and 'Limitations of using WER alone'. These sections highlight crucial real-world considerations for a research-oriented engineer.

Here are the key points to remember:

  • WER can exceed 100%: If the model's output is much longer than the reference (i.e., has many insertions), the numerator can be larger than the denominator . This indicates a very poor transcription.
  • Semantic Blindness: WER treats all word errors equally. However, a single word substitution can completely change the meaning of a sentence.
    • turn left -> turn right (1 substitution, critical error)
    • turn uh left -> turn left (1 deletion, minor error)
      Both could result in a similar WER, but the semantic impact is vastly different.
  • Context is Key: What constitutes a "good" WER is domain-specific. A WER of 15% might be acceptable for casual voice commands, but for medical transcription, the target might be below 3%.
  • Analyze Error Types: Don't just look at the final WER. Analyzing the breakdown of S, D, and I can give you clues about your model's failure modes. A high number of insertions might point to issues with silence detection, while a high number of substitutions could indicate acoustic ambiguity.

Conclusion

In this lesson, we have thoroughly covered the standard metrics for evaluating ASR systems. This knowledge is not just academic; it is the fundamental way in which progress in the field is measured and reported in research papers and production systems.

Key Takeaways:

  • WER and CER are the standard metrics for ASR, quantifying the "distance" between a model's hypothesis and a ground-truth reference.
  • They are calculated based on the Levenshtein distance, counting the minimum number of Substitutions, Deletions, and Insertions needed to transform one sequence into another.
  • The formula is , where is the length of the reference.
  • Libraries like jiwer make it easy to implement these metrics in Python.
  • Interpretation is crucial: WER is semantically blind, and a "good" score is highly dependent on the specific application.

With this lesson, we conclude our foundational module on audio data augmentation and pipelines. You are now equipped with the skills to prepare, process, and evaluate data for sophisticated speech AI tasks.

Preview of the Next Lesson:
We are now ready to dive into the models themselves. Our next module, "Sequence Modeling with Transformers," will start by exploring how 1D Convolutional Neural Networks (CNNs) can act as feature extractors for raw audio waveforms. This will be our first step in building deep learning models that can learn directly from sound.

Can't find a good explanation? Sign up and we'll make it for you

Sign up