Hello! Welcome to your next lesson in our exploration of Audio Language Models.
In our previous lesson, we established a crucial foundation: how neural audio codecs like SoundStream and EnCodec can take a continuous audio waveform and "tokenize" it into a sequence of discrete integer codes. This process effectively translates audio into a format that a language model can understand.
Today, we will build directly on that foundation to address the learning outcome: Describe how a standard language model architecture (e.g., a decoder-only Transformer) can be applied to discrete audio tokens for generation (AudioLM).
We will see how the language modeling paradigm, which has revolutionized natural language processing, can be ported directly to the audio domain. This is the core concept behind Audio Language Models (AudioLMs) and is a critical step on our journey toward building models that can "listen" and "talk."
1. The Paradigm Shift: Treating Audio as a Language
The central insight is this: if we can represent audio as a sequence of discrete tokens, then generating audio becomes a language modeling task. The problem shifts from predicting continuous sample values to predicting the next discrete token in a sequence.
audio_waveform → [codec_encoder] → [token_1, token_2, ..., token_N]
[token_1, ..., token_k] → [Language_Model] → predict token_{k+1}
This allows us to leverage the immense power of autoregressive Transformer models, like those in the GPT family, which excel at next-token prediction. By compressing audio, we dramatically shorten the sequence length, allowing the model's self-attention mechanism to capture much longer-term dependencies than would be possible with raw audio.
Let's explore how this is done in practice.
Neural audio codecs: how to get audio into LLMs
The article 'Neural audio codecs: how to get audio into LLMs' from Kyutai provides an excellent, hands-on walkthrough of this exact process. It demonstrates training a GPT-style model on audio tokens.
Please read the following sections: 'Why care about audio': This briefly recaps the motivation for creating audio generation models. 'Dealing with multiple levels': This is a key practical step. It explains how the multiple code streams from a Residual Vector Quantizer (RVQ) are flattened into a single sequence for the LLM. 'Finally, let's train': This section shows that once the data is tokenized and flattened, training is identical to training a text LLM. Notice the results from their simple, custom codec. 'How far can a codec get us?': This demonstrates the dramatic improvement in generation quality when using a more advanced codec (Mimi) compared to the simple one. Listen to the audio samples to appreciate the difference.
As the article shows, the quality of the underlying codec is paramount. A better tokenizer leads to better generation. However, even with a great codec, just flattening the RVQ tokens and training a single language model has its limits. A more sophisticated approach is needed to achieve both long-term coherence and high acoustic fidelity. This brings us to AudioLM.
2. AudioLM: A Hierarchical Language Model for Audio
AudioLM, developed by researchers at Google, was a landmark paper that truly systematized the "audio as language" approach. Instead of treating all audio tokens equally, it introduced a hierarchical modeling strategy that separates the what from the how.
To get a high-level intuition, let's hear from one of SoundStream's authors as he introduces the generative modeling work that followed it.
Neil Zeghidour: SoundStream: an end-to-end neural audio codec
In this final segment of his talk on SoundStream, Neil Zeghidour pivots to discuss how the discrete tokens produced by the codec are used in the generative model, AudioLM.
Watch the section titled 'Follow-up: generative models' (36:55 - 45:06). Focus on: The core idea of treating audio generation as language modeling on discrete SoundStream tokens. The challenge of flattening the sequence from the residual quantizer, which multiplies the sequence length. The impressive quality and coherence of the speech and music continuation demos. Notice how the model preserves speaker identity, accent, and acoustic conditions.
The results are striking. The model doesn't just babble; it generates coherent speech and music that maintains context. How does AudioLM achieve this? Through two key innovations: hybrid tokenization and hierarchical generation.
2.1. Hybrid Tokenization: Semantic and Acoustic Tokens
AudioLM uses two parallel streams of tokens derived from the same input audio:
-
Acoustic Tokens: These are generated by a neural audio codec, specifically SoundStream. As we know from the previous lesson, these tokens (from the RVQ) are designed to reconstruct the waveform with high fidelity. They capture the low-level acoustic properties: timbre, pitch, reverberation, and speaker identity.
-
Semantic Tokens: These are generated from a different model, typically a large, self-supervised speech model like w2v-BERT or HuBERT. The intermediate representations from this model are clustered (using k-means) to create a discrete set of "semantic" tokens. These tokens capture higher-level linguistic information—the content of what's being said—while being largely invariant to the specific acoustic details.

2.2. Hierarchical Autoregressive Modeling
With these two token types, AudioLM employs a multi-stage generative process, using a separate decoder-only Transformer at each stage. The generation is factored hierarchically: first, model the semantics, then, conditioned on the semantics, model the acoustics.
Mathematically, the joint probability of the full token sequence , composed of semantic tokens and acoustic tokens , is factored as:
This factorization is implemented with a cascade of Transformer models.

The process, simplified, is as follows:
-
Semantic Generation: A Transformer LM is trained to predict the next semantic token, , given the previous ones, . During inference, it generates the entire sequence of semantic tokens, establishing the high-level content and structure of the audio to be generated.
-
Acoustic Generation: A second, larger Transformer LM is then conditioned on the full sequence of generated semantic tokens. It autoregressively generates the acoustic tokens from the SoundStream codec, , given all previous acoustic tokens and the semantic tokens . This stage effectively "renders" the semantic plan into high-fidelity sound.
As the diagram shows, the acoustic generation stage can be further broken down into "coarse" and "fine" stages, where the first few RVQ levels are generated first, followed by the later levels, allowing the model to focus on different levels of acoustic detail sequentially.
For a more detailed technical breakdown, the following resource is excellent.
AudioLM: Token-Based Audio Generation - Emergent Mind
This summary from Emergent Mind provides a concise, academic overview of the AudioLM architecture.
Read sections 1 through 4. Focus on: Section 2 (Neural Audio Tokenization Pipeline): Reinforces the distinction between semantic (w2v-BERT) and acoustic (SoundStream) tokens. Section 3 (Hierarchical Language Modeling Architecture): This is the core of the lesson. Understand the probabilistic factorization and how each stage uses a decoder-only Transformer. Section 4 (Training and Inference Procedures): Note that each stage is trained independently using a standard next-token prediction objective (cross-entropy loss). Then, see how they are chained together for generation.
3. The Big Picture: Training and Generation
Despite the complex, multi-stage architecture, the training of each individual component in AudioLM is remarkably simple. Each of the Transformer language models—the one for semantic tokens and the one(s) for acoustic tokens—is trained independently on its respective token sequence using a standard autoregressive next-token prediction objective with a cross-entropy loss.
The magic happens during inference, where the models are chained together:
- Provide a short audio prompt.
- Run the prompt through the tokenization pipeline to get initial semantic and acoustic tokens.
- The semantic LM autoregressively generates a continuation of the semantic tokens.
- The acoustic LM takes the full sequence of semantic tokens (prompt + generated) and the prompt's acoustic tokens, and autoregressively generates the corresponding acoustic tokens.
- The complete stream of generated acoustic tokens is passed to the SoundStream decoder to synthesize the final audio waveform.
This hierarchical approach allows AudioLM to generate audio that is not only locally high-quality (thanks to the SoundStream codec) but also globally coherent and semantically meaningful over long durations (thanks to the semantic tokens and hierarchical modeling).
Conclusion
In this lesson, we have seen how the principles of language modeling can be powerfully applied to audio generation. By discretizing audio into tokens, we can use standard Transformer architectures to create sophisticated generative models.
Key Takeaways:
- Audio as Language: By converting audio into discrete tokens via a neural codec, audio generation can be framed as a next-token prediction task, just like text generation.
- Standard LM Application: A decoder-only Transformer can be trained on a flattened sequence of audio tokens using an autoregressive cross-entropy loss. The model learns to predict the next audio token in the sequence.
- AudioLM's Innovation: AudioLM refines this approach with a hierarchical structure. It separates audio information into semantic tokens (for content and structure) and acoustic tokens (for sound quality and speaker identity).
- Hierarchical Generation: AudioLM uses a cascade of Transformer models. A first model generates the semantic "blueprint," and subsequent models generate the acoustic details conditioned on that blueprint. This enables both long-term coherence and high acoustic fidelity.
Preview of the Next Lesson:
We've seen how AudioLM can generate coherent continuations of an audio prompt. But what if we want more explicit control? The next generation of models, like VALL-E, leverages the same token-based paradigm to perform "in-context learning" for audio. In the next lesson, we will explore the mechanics of in-context learning for audio generation, where a model can clone a voice from a short 3-second sample and use it to synthesize new speech from a text prompt.