Hello! Welcome to the next lesson in our module on model optimization and deployment.
In our last session, we saw how exporting a model to ONNX and running it with ONNX Runtime can significantly reduce inference latency. You benchmarked the performance gain, turning an abstract optimization into a concrete speedup. This low latency is not just an academic exercise; it's a critical enabler for real-time applications.
Today, we'll explore one of the most demanding of these applications. Our learning outcome is to describe the architectural requirements for streaming ASR systems and their latency-accuracy trade-offs. We will move beyond processing pre-recorded files (batch ASR) and dive into the world of transcribing audio as it is being spoken (streaming ASR). You will learn what makes a model "streamable," the specific architectural designs required, and the inherent compromises between speed and accuracy.
1. Batch vs. Streaming ASR: Defining the Problem
First, let's establish a clear distinction between the two primary modes of speech recognition.
- Batch ASR: The model receives a complete audio file, processes the entire thing, and returns the full transcription. This is what you've worked with so far when using models like Whisper on a
.wavfile. The model has access to the full context of the utterance, which often leads to higher accuracy. - Streaming ASR: The model receives a continuous stream of audio data in small chunks and must produce a transcription in near real-time, with a delay of no more than a few hundred milliseconds to a couple of seconds. This is essential for applications like live captioning, voice assistants, and real-time call transcription.
The following video provides an excellent introduction to this concept and the inherent trade-offs.
Can Whisper be used for real-time streaming ASR?
Watch this segment from 'Can Whisper be used for real-time streaming ASR?' by the Efficient NLP channel to understand the fundamental difference between batch and streaming ASR.
Watch from 00:33 to 01:19. The video clearly defines streaming ASR and points out the core trade-off: streaming models are generally expected to have slightly lower accuracy than batch models because they lack full future context.
As the video highlights, the core challenge of streaming ASR is operating under two strict constraints, which are fundamentally at odds with how models like the Transformer were originally designed:
- Causality: The model cannot use future audio frames that have not yet arrived. It must make predictions based only on past and present context.
- Low Latency: The model has a very small budget for "lookahead." It can only buffer a tiny amount of audio (e.g., a few hundred milliseconds) before it must start producing output.
This forces a critical trade-off between acoustic evidence, language model predictions, and the delay a user experiences.
2. Architectural Requirements for Streaming
You can't just take any ASR model and "make it stream." As you'll see, attempting to do so with a batch model like Whisper involves inefficient heuristics. A true streaming model must be designed and trained with streaming in mind from the ground up.
Let's explore the fundamental architectural modifications required to meet the causality and low-latency constraints. The SpeechBrain tutorial on streaming Conformers is an excellent, in-depth resource for this.
Streaming Speech Recognition with Conformers
The following sections from the SpeechBrain tutorial, 'Streaming Speech Recognition with Conformers', break down the core engineering challenges and architectural solutions for building a streaming model. We will read it in parts.
First, read the section 'What a streaming model needs to achieve'. This section sets the stage by highlighting the need to restrict context and the different performance characteristics between training and inference.
As the text explains, the key is to manage context. Let's break down how this is achieved in different parts of a modern ASR architecture like the Conformer.
Requirement 1: Chunk-Based Attention
The self-attention mechanism in a Transformer is non-causal by nature; every token attends to every other token in the sequence. This is incompatible with streaming. While a simple solution is causal attention (where a token can only attend to past tokens), a more effective approach for ASR is Chunked Attention.

Now, let's dive into the specifics of how this works.
Streaming Speech Recognition with Conformers
Let's continue with the SpeechBrain tutorial to understand how Chunked Attention works.
Read the section 'What is Chunked Attention, and why do we prefer it?'. Focus on understanding that frames within a chunk can attend to each other, and chunks can attend to a limited number of past chunks. This is enforced during training with a special attention mask.
During inference, this chunking strategy allows the model to process incoming audio piece by piece. The "state" of the conversation (i.e., the context from previous chunks) is managed by caching.
Requirement 2: Efficient State Management (Caching)
A naive streaming implementation might re-process the left context repeatedly for every new chunk. This is incredibly inefficient. A proper streaming architecture caches the necessary hidden states from previous chunks to be reused for the next one.
For Transformer-based models, this typically involves a KV Cache. The Keys (K) and Values (V) from the self-attention layers of the left-context chunks are stored. When a new chunk arrives, the model computes its new Queries (Q) and attends to both its own K/V pairs and the cached K/V pairs from the left context.
A crucial technique for very long audio streams is KV Cache Eviction. Since the model only needs to look at a recent window of audio (e.g., the last minute), the parts of the KV cache corresponding to very old audio can be discarded ("evicted") to prevent GPU memory from growing indefinitely.
Inference Characteristics of Streaming Speech Recognition
This video, 'Inference Characteristics of Streaming Speech Recognition', discusses the practical aspects of deploying streaming models and explains the concept of KV cache eviction.
Watch from 07:22 to 08:18. Pay close attention to how the video contrasts standard LLM inference (where the KV cache grows with the sequence) with streaming ASR, where old parts of the cache can be evicted to enable indefinite streaming.
Requirement 3: Streaming-Aware Convolutions
The Conformer architecture, which you've encountered before, heavily uses convolution modules. Standard convolutions also have a lookahead dependency (a "future context"), which breaks causality.
The solution is analogous to chunked attention: Dynamic Chunk Convolutions. The convolution's receptive field is masked so that it cannot see input frames that belong to a future chunk.
Streaming Speech Recognition with Conformers
Let's return to the SpeechBrain tutorial to see how convolutions are handled.
Read the section 'Dynamic Chunk Convolutions'. The key takeaway is that the convolution is modified to respect the same chunk boundaries used by the attention mechanism, preventing it from depending on future, unseen audio.
Requirement 4: Dynamic Chunk Training
If a model is trained with a fixed chunk size (e.g., 640ms), it will perform poorly at inference if a different chunk size is used. This forces a fixed latency-accuracy trade-off.
A powerful technique to overcome this is Dynamic Chunk Training (DCT). During training, for each batch, a random chunk size and left-context size are sampled. This trains a single, robust model that can be deployed with any chunk size at inference time, allowing the developer to choose the desired latency-accuracy trade-off on the fly.
Streaming Speech Recognition with Conformers
Finally, let's look at the training strategy from the SpeechBrain tutorial.
Read the section 'How to pick the chunk size?'. Understand the DCT strategy: training on a mix of full-context utterances and randomly-sized chunks. This results in a single, flexible model for both streaming and non-streaming inference.
3. A Comparison of Streaming ASR Architectures
Now that we understand the core requirements, let's compare the three main families of streaming ASR architectures. Each meets the streaming constraints differently, leading to distinct behaviors and trade-offs.
The following article provides a superb deep dive into how these architectures work and why their design choices matter for the user experience.
Voice AI Deep Dive: Streaming ASR Architectures (CTC vs RNN-T ...
This blog post, 'Voice AI Deep Dive: Streaming ASR Architectures', offers a clear comparison of the dominant streaming ASR model families.
Please read the sections on CTC, RNN-T, and Attention/Seq2Seq. For each one, focus on: The core intuition of how it works. Why it is (or isn't) friendly to streaming. Its primary strengths and weaknesses.
To summarize the key points from the article:
-
Connectionist Temporal Classification (CTC):
- How it works: Introduces a special
blanktoken and allows for monotonic alignment between audio frames and output labels. The model emits spikes of non-blank labels only when it has strong acoustic evidence. You will derive the CTC loss in a future lesson. - Streaming-friendliness: Excellent. The frame-by-frame processing and monotonic alignment are naturally suited for low-latency output.
- Trade-off: The language modeling is implicit and weak. CTC models rely heavily on the acoustic encoder and often need an external language model fused during decoding to achieve high accuracy, especially for ambiguous phrases.
- How it works: Introduces a special
-
Recurrent Neural Network Transducer (RNN-T):
- How it works: An evolution of CTC. It adds a separate Prediction Network that acts like an internal language model, consuming the previously generated text tokens. A Joint Network then combines the acoustic information from the encoder and the linguistic context from the prediction network to make a final decision.
- Streaming-friendliness: Excellent. It maintains the monotonic alignment of CTC while integrating a powerful language model directly into the architecture. This is why it dominates production streaming systems (e.g., in Google's voice products).
-
Chunked Attention / Seq2Seq:
- How it works: This is the approach we studied in detail with the Conformer model. It takes powerful offline encoder-decoder models and adapts them for streaming using the chunking and caching mechanisms discussed earlier.
- Streaming-friendliness: Good, but requires significant "surgery." It's an adaptation, not a native design.
- Trade-off: Can achieve very high accuracy and long-context consistency but is susceptible to boundary artifacts (e.g., words being split or revised at chunk edges). The latency is directly tied to the chosen chunk size.
4. The Latency-Accuracy Trade-off Matrix
The choice of architecture and its configuration parameters (like chunk size) directly impacts the user experience. The article you just read provides a fantastic summary table comparing these models on practical metrics.
Voice AI Deep Dive: Streaming ASR Architectures (CTC vs RNN-T ...
Let's revisit the 'Voice AI Deep Dive' article for its summary table and conclusion.
Review the 'Practical Comparison Matrix' and the 'What to Choose' recommendation section. This synthesizes the trade-offs into a clear, actionable guide.
Here is the comparison matrix from the article, which is a crucial summary of this lesson's topic:
| Property | CTC | RNN-T | Chunked Attention |
|---|---|---|---|
| TTFT | Excellent | Excellent | Good (chunk-dependent) |
| Partial stability | Good | Good–Excellent | Variable |
| Final-word latency | Good | Excellent | Variable (often higher) |
| LM integration | External or implicit | Built-in | Built-in |
| Streaming complexity | Low–Medium | Medium | Medium–High |
| Boundary artifacts | Low | Low | Medium–High |
(TTFT = Time-to-first-token)
As the Dynamic Chunk Training strategy showed, even within a single architecture like Chunked Attention, you can tune the trade-off. This is most directly controlled by chunk size.
Streaming Speech Recognition with Conformers
Finally, let's look at a concrete example of the chunk size vs. accuracy trade-off from the SpeechBrain tutorial.
Review the section 'What metrics does the chunk size and left context size impact?'. Observe the graph showing WER vs. chunk size. It clearly demonstrates that as chunk size decreases (improving latency), the Word Error Rate (WER) tends to increase (degrading accuracy).
Conclusion
In this lesson, you have moved from the general concept of low-latency inference to the specific, rigorous architectural demands of streaming ASR. You now understand that building a real-time speech recognition system is not just about having a fast model, but about having the right architecture for the task.
Key Takeaways:
- Streaming ASR operates under causality and low-latency constraints, fundamentally differing from batch ASR.
- True streaming models require specific architectural designs, not just post-hoc adaptations. Key requirements include chunked processing (like Chunked Attention and Dynamic Chunk Convolutions) and efficient state management (caching and KV cache eviction).
- Dynamic Chunk Training is a key strategy that produces a single flexible model capable of operating at different latency-accuracy points.
- The dominant streaming architectures are CTC, RNN-T, and Chunked Attention, each with distinct trade-offs. RNN-T is often preferred in production for its strong balance of accuracy and speed due to its integrated language model.
- The primary latency-accuracy trade-off is controlled by chunk size: smaller chunks yield lower latency but can increase error rates.
Preview of the Next Lesson:
Having covered the theory, our next step is to put it into practice. In the next lesson, you will implement a basic streaming inference loop for a frame-synchronous ASR model. You will write code that simulates receiving audio chunk-by-chunk and uses a streaming-capable model to perform transcription incrementally.