Skip to main content
Create your own
Lesson illustration

Audio-Text LLM Architecture: AudioPaLM Case Study

Hello! Welcome back to our course.

In the last lesson, we got our hands dirty with practical speech-to-speech translation, running recipes in major toolkits like Fairseq and ESPnet. We saw how these specialized, often complex pipelines are constructed to handle a single primary task.

Today, we shift our focus to a more recent and powerful paradigm: large language models that are inherently multimodal, capable of understanding and generating both text and audio within a single, unified architecture. This approach aims to leverage the vast world knowledge and reasoning capabilities of text-based LLMs for sophisticated audio tasks.

Your learning outcome for this lesson is to describe the architecture of multimodal audio-text LLMs like AudioPaLM. We will explore how these models are built, trained, and why they represent a significant step forward from the systems we've studied so far.

1. From Cascaded Systems to Integrated Speech LLMs

Before we dive into a specific architecture, let's understand the motivation behind this new approach. A common way to build a conversational AI is to "cascade" models: an ASR model converts speech to text, a text-based LLM processes the text, and a TTS model converts the response back to speech. While functional, this approach has several drawbacks.

To get a clear overview of these issues and the alternative "Speech LLM" paradigm, please watch the beginning of the following video.

Speech LLMs: Models that listen and talk back

This video from Efficient NLP concisely explains why integrated, end-to-end models are superior to traditional cascaded systems.

Watch the first two minutes of the video (00:00 - 01:55). Pay close attention to the three main problems with cascading models that are discussed.

As the video explains, cascading models suffer from:

  1. Loss of Information: Critical paralinguistic information like tone, emotion, and speaker identity is lost when speech is converted to plain text.
  2. Error Propagation: An error made by the ASR model (e.g., mishearing "peaches" as "beaches") is passed down to the LLM, which has no way to correct it, potentially leading to nonsensical responses.
  3. Increased Latency: Each model in the chain must wait for the previous one to finish, making real-time conversation difficult.

Integrated audio-text LLMs aim to solve these problems by processing audio more directly, preserving the rich information it contains.

2. The Foundation: Treating Audio as a Language

For a text-based LLM to process audio, the continuous audio waveform must be converted into a sequence of discrete tokens, just like words in a sentence. This process of discrete audio tokenization is the cornerstone of modern audio language models.

There are two main philosophies for creating these tokens, which capture different aspects of the audio signal.

Discrete Audio Token Generation - Emergent Mind

This article from Emergent Mind provides an excellent technical summary of discrete audio token generation. It clearly defines the different types of tokens that are fundamental to models like AudioPaLM.

Read the sections '1. Motivations and Foundations' and '2. Tokenization Methodologies'. Focus on understanding the distinction between 'Acoustic tokens' and 'Semantic tokens'.

As the article details:

  • Acoustic Tokens: These are generated by neural audio codecs like EnCodec or SoundStream, which we've encountered before. They aim to capture the low-level acoustic details of the waveform, enabling high-fidelity reconstruction. They are good for sound, but poor for meaning.
  • Semantic Tokens: These are extracted from the hidden states of self-supervised learning (SSL) models like HuBERT or w2v-BERT. They capture higher-level linguistic content (phonemes, word-like units) but discard most of the acoustic detail. They are good for meaning, but poor for reconstructing the original sound.

The key insight of audio language models is that if we can represent audio as a sequence of these tokens, we can apply the same powerful Transformer architectures used for natural language processing.

3. Architectural Deep Dive: AudioPaLM

AudioPaLM is a landmark model from Google that exemplifies this new architecture. It fuses a powerful, pre-trained text LLM (PaLM-2) with a sophisticated audio generation model (AudioLM) to create a single system that can "speak" and "listen."

Let's dissect its architecture piece by piece, using the original paper and a helpful diagram.

AudioPaLM Model Architecture for Multimodal Tasks
This diagram illustrates the unified architecture of AudioPaLM. It shows how both audio and text inputs are tokenized and processed by a single decoder-only Transformer, which can then generate either text or audio tokens as output.

3.1. The Core Principle: A Unified Vocabulary

The central idea of AudioPaLM is to create a single, unified vocabulary that contains tokens for both text and audio.

AudioPaLM: A Large Language Model That Can Speak and Listen

Let's turn to the official AudioPaLM paper. This section describes how the model represents both audio and text in a way that a Transformer can understand.

Read Section 3.1, 'Audio Embeddings and Tokenization'. Note the models mentioned (w2v-BERT, USM) for creating semantic audio tokens.

As described in the paper and shown in the diagram, the process is:

  1. Text Tokenization: Standard text is tokenized into text tokens using a SentencePiece model.
  2. Audio Tokenization: An input audio waveform is processed by a pre-trained speech model (like USM, the Universal Speech Model) to extract semantic tokens. These tokens represent the linguistic content of the speech.
  3. Joint Vocabulary: These two sets of tokens are combined into a single, larger vocabulary. For the model, there's no fundamental difference between a text token and an audio token; they are just integers from a shared vocabulary.

3.2. The LLM Core: Modifying a Text LLM

With a unified vocabulary, how do we adapt a text-only LLM like PaLM-2 to handle audio? The modification is surprisingly minimal and elegant.

AudioPaLM: A Large Language Model That Can Speak and Listen

This next section from the AudioPaLM paper reveals the key architectural change made to the PaLM-2 model.

Read Section 3.2, 'Modifying text-only decoders to model both text and audio'. Focus on understanding what the 'token embeddings matrix' is and how it's expanded.

The only change required is to the token embedding matrix.

  • A standard decoder-only Transformer has an embedding matrix of size , where is the text vocabulary size and is the embedding dimension. This matrix maps each token ID to a dense vector.
  • To create AudioPaLM, this matrix is simply expanded to size , where is the size of the audio token vocabulary.
  • The first rows (for text tokens) are initialized with the weights from the pre-trained PaLM-2 model.
  • The new rows (for audio tokens) are randomly initialized.

The rest of the massive Transformer architecture remains completely untouched. It continues to be a decoder-only model that predicts the next token in a sequence, blissfully unaware of whether the tokens represent English text, French speech, or a mix of both.

3.3. Training and Generation

The model is then fine-tuned on a diverse mixture of tasks. The desired task is specified using a simple text prefix. For example:

  • Input: [ASR French] + (French audio tokens) -> Output: (French text tokens)
  • Input: [S2ST English French] + (English audio tokens) -> Output: (French audio tokens)

This multi-task, multi-modal training allows the model to learn the relationships between speech and text across different languages.

When generating speech, AudioPaLM follows the AudioLM procedure we've discussed previously:

  1. The main Transformer autoregressively generates a sequence of semantic audio tokens.
  2. These semantic tokens are passed to a separate decoder (in AudioPaLM's case, a non-autoregressive model called SoundStorm) which generates the fine-grained acoustic tokens.
  3. A final vocoder synthesizes the waveform from the acoustic tokens.

This hierarchical generation process allows the model to preserve paralinguistic features like speaker identity, even during translation, because the semantic tokens from the LLM guide the generation without strictly dictating the acoustic properties.

3.4. The Power of Pre-training

A crucial finding of the AudioPaLM paper is the immense benefit of starting from a pre-trained text LLM.

AudioPaLM: A Large Language Model That Can Speak and Listen

This ablation study from the paper provides clear evidence for the effectiveness of the fine-tuning approach.

Quickly review Sections 5.4.2 ('Training from scratch vs. finetuning') and 5.4.8 ('Impact of using PaLM-2'). Compare the performance numbers in Table 6 and Table 12. You don't need to analyze them deeply, just observe the trend.

The results are stark:

  • Finetuning vs. From Scratch (Table 6): The model fine-tuned from a PaLM checkpoint dramatically outperforms an identical model trained from scratch on the same speech data. This proves that the LLM's pre-existing linguistic knowledge is successfully transferred to speech tasks.
  • PaLM vs. PaLM-2 (Table 12): Using a better text LLM (PaLM-2) as the base generally leads to better performance on speech tasks, especially speech translation. This shows that improvements in the text domain directly benefit the audio domain.

4. An Alternative Approach: Qwen-Audio

AudioPaLM's approach of integrating at the token vocabulary level is powerful, but not the only way. Let's look at another model, Qwen-Audio, to see a different integration strategy.

Multi modal Audio + Text Fine tuning and Inference with Qwen

This video from Trelis Research provides a great walkthrough of the Qwen-Audio architecture. It serves as an excellent comparison to AudioPaLM.

Watch the section from 01:19 to 06:49. Focus on: The components used: a Whisper model as the audio encoder and a Qwen model as the LLM. How the outputs from the audio encoder are adapted to be used by the LLM (the 'linear layer'). The presenter's explanation of why this integrated model is better than a simple cascaded pipeline.

Qwen-Audio's architecture is different from AudioPaLM's in a key way:

  • AudioPaLM: Integrates at the token level. Audio is turned into discrete tokens that share a vocabulary with text tokens.
  • Qwen-Audio: Integrates at the feature level.
    1. An audio encoder (Whisper's encoder) processes the audio and outputs a sequence of continuous feature vectors (embeddings).
    2. A simple linear projection layer transforms these audio feature vectors into the same vector space used by the Qwen LLM's text embeddings.
    3. The LLM then receives a sequence that is a concatenation of text embeddings and the projected audio embeddings.

This is more akin to connecting two large pre-trained models with a small "adaptor" module, rather than modifying the vocabulary of one model. Both are valid and effective strategies for building powerful audio-text LLMs.

Conclusion

In this lesson, we explored the architecture of modern multimodal audio-text LLMs. You've seen that the core innovation is to find a way for powerful, pre-trained text language models to process audio, thereby inheriting their vast linguistic knowledge.

Key Takeaways:

  • Multimodal audio-text LLMs like AudioPaLM and Qwen-Audio overcome the limitations of older cascaded systems (information loss, error propagation, latency).
  • The foundation of these models is discrete audio tokenization, which converts continuous waveforms into sequences that a Transformer can process.
  • AudioPaLM's architecture works by creating a unified vocabulary for text and audio tokens and expanding the embedding matrix of a pre-trained text LLM (PaLM-2) to accommodate them.
  • Qwen-Audio presents an alternative architecture, using a linear adaptor to project features from a speech encoder (Whisper) into the LLM's embedding space.
  • The most effective training strategy is to fine-tune a powerful, pre-trained text LLM on a mixture of speech-text tasks, rather than training a multimodal model from scratch.

Preview of the Next Lesson:

While these models are incredibly powerful, they are not without their limitations. In our next and final lesson of this module, we will discuss the challenges and future directions in audio language modeling, including issues with long-form generation, controlling expressiveness, and the ongoing quest for the perfect audio tokenizer.

Can't find a good explanation? Sign up and we'll make it for you

Sign up