Skip to main content
Create your own

Beyond LoRA: Exploring Other PEFT Methods

Hello! In our previous lesson, we explored LoRA, a powerful technique that makes fine-tuning large models manageable by modifying a tiny fraction of their internal weights. This is an example of a reparameterization method, where we approximate the weight update with low-rank matrices.

Today, we'll shift our focus to a different family of Parameter-Efficient Fine-Tuning (PEFT) methods. Instead of changing the model's weights, these techniques add new, trainable parameters to the input or hidden states. Your learning outcome for this lesson is to apply other PEFT techniques like prefix and prompt tuning. We'll investigate how these methods work, their trade-offs, and where they fit in the broader PEFT landscape.

1. Hard Prompts vs. Soft Prompts: A New Paradigm

You're already familiar with prompt engineering, where you manually craft a text prompt to guide a model's output. This is often called using hard prompts or discrete prompts. It's an art that involves finding the right words in the token space.

PEFT introduces a new concept: soft prompts. Instead of searching for the perfect text prompt, we let the model learn the perfect prompt for a task directly in the embedding space. A soft prompt is a sequence of trainable vectors (or "virtual tokens") that don't correspond to any real words but are optimized via backpropagation to steer the model's behavior.

The following videos provide an excellent introduction to this idea, contrasting manual prompt engineering with the automated learning of soft prompts in prompt tuning.

What is Prompt Tuning?

First, this video from IBM Technology gives a high-level overview, clearly distinguishing between fine-tuning, prompt engineering (hard prompts), and prompt tuning (soft prompts).

Watch from 00:45 to 07:20. Focus on: The definition of prompt tuning and its goal: tailoring a model to a narrow task with limited data. The distinction between human-written 'hard prompts' and AI-generated 'soft prompts' (uninterpretable strings of numbers). The visual summary at the end that compares how fine-tuning, prompt engineering, and prompt tuning interact with a pre-trained model.

LLM2 Module 2 - Efficient Fine-Tuning | 2.3 PEFT and Soft Prompt

Next, this video from Databricks offers a more technical look at soft prompts and how they are integrated into the model's input.

Watch the first 7 minutes and 57 seconds (00:00 - 07:57). Pay attention to: The three categories of PEFT: additive, selective, and reparameterization (LoRA is reparameterization). The concept of 'virtual tokens' that are concatenated with the input text embeddings. How these virtual tokens are learned through backpropagation while the foundation model's weights are kept frozen.

2. Prompt Tuning: Adding Trainable Tokens to the Input

As you just saw, prompt tuning is the most direct application of soft prompts. It's an additive method that involves:

  1. Freezing the entire pre-trained LLM. Not a single one of its billions of parameters will be updated.
  2. Prepending a short sequence of trainable virtual token embeddings to the input sequence.
  3. Training only these virtual tokens. The model learns, through backpropagation, the optimal embedding values for these tokens to solve the specific downstream task.
Model Tuning vs. Prompt Tuning Comparison
This diagram from the original Prompt Tuning paper starkly contrasts full model tuning with prompt tuning. Instead of creating a full 11B-parameter copy of a model for each task, you only create and store a tiny (e.g., 20K parameter) task-specific prompt.

Key Characteristics:

  • Extreme Parameter Efficiency: Prompt tuning updates the fewest parameters of almost any PEFT method. For a model like GPT-3, you might train only a few hundred thousand parameters instead of 175 billion.
  • Multi-Task Serving: Because the base model is shared and frozen, you can serve many different tasks using a single LLM in memory. To switch tasks, you simply swap out the tiny soft prompt file—a huge operational advantage.
  • Performance: A key finding is that prompt tuning becomes competitive with full fine-tuning only on very large models (e.g., >10B parameters). On smaller models, it often underperforms LoRA or full fine-tuning.
  • Interpretability: The learned virtual tokens are just vectors of numbers and are not human-interpretable.

3. Prefix Tuning: A More Powerful Approach

Prefix tuning is a more powerful evolution of prompt tuning. While prompt tuning adds trainable tokens only at the input embedding layer, prefix tuning injects trainable parameters at every layer of the Transformer.

Here's how it works:

  1. The base model is frozen.
  2. For each Transformer layer, a small, trainable "prefix" matrix is created.
  3. This prefix is concatenated to the Key () and Value () matrices within the self-attention mechanism for that layer.
  4. Only these prefix matrices across all layers are trained.

By influencing the attention calculation at every layer, the model can learn a much more nuanced, task-specific behavior.

Comparison of Fine-tuning and Prefix-tuning for Transformers
This visual shows that instead of training a whole new model for each task, prefix-tuning uses a single pretrained model and simply adds a small, task-specific prefix that influences the model's internal computations.

Prompt Tuning vs. Prefix Tuning

Feature Prompt Tuning Prefix Tuning
Location of Added Params Only at the input embedding layer. At every Transformer layer (in the attention block).
Number of Trainable Params Extremely low. Very low, but more than Prompt Tuning.
Expressiveness Less expressive. Influences the model only at the start. More expressive. Guides the model's computation at every depth.
Performance Good on huge models, weaker on smaller ones. Generally outperforms Prompt Tuning, especially on generation tasks.

This article from the Hugging Face blog provides a good summary of these concepts.

PEFT: Parameter-Efficient Fine-Tuning Methods for LLMs

To solidify your understanding, please read the following sections from this Hugging Face blog post on PEFT. It cleanly summarizes and contrasts Prompt Tuning and Prefix Tuning.

Read the sections titled 'Prompt Tuning' and 'Prefix Tuning' under the 'Soft Prompts' heading. Focus on the conceptual differences in their implementation and performance characteristics.

Test your understanding!

Imagine you are trying to fine-tune a model for a highly complex code generation task. The model needs to understand and maintain context over very long sequences and adhere to strict syntactical rules. Based on what you've learned, would you choose Prompt Tuning or Prefix Tuning? Why?

Show answer

Prefix Tuning would be the more suitable choice.

Reasoning:
Code generation is a complex task that requires nuanced control at various levels of abstraction—from high-level logic to low-level syntax.

  • Prompt Tuning only provides an initial "nudge" to the model at the input layer. It might set a general direction, but it can't guide the layer-by-layer processing required for a complex generative task.
  • Prefix Tuning, by injecting a trainable prefix into every layer, provides a continuous, task-specific context throughout the entire computational process. This allows the model to adapt its internal representations and attention patterns at each level of depth, which is crucial for maintaining logical consistency and syntactical correctness in code generation.

4. The Broader PEFT Landscape

Prompt and Prefix Tuning are just two of many PEFT techniques. Your goal is a comprehensive understanding, so let's briefly survey the landscape to see how these methods relate to others. LoRA, which you've already studied, is a reparameterization method. The methods below are primarily additive or selective.

The following article gives an excellent overview and benchmark of many popular PEFT methods.

Tuning LLM Efficiently: A Deep Dive into PEFT

This article from Dataroot Labs provides a fantastic survey of various PEFT methods and, most importantly, benchmarks their performance against each other.

Please read the following sections: 'Prompt tuning vs Prefix tuning': A quick review of what we've just covered. 'Adapters': The original additive method that inspired many others. 'LoRA': A brief recap of what you learned last lesson. 'Selective PEFT' ('BitFit' and 'LayerNorm Tuning'): Understand this different philosophy of tuning a small subset of existing parameters rather than adding new ones. 'Benchmarking PEFT Strategies' and 'Conclusion': This is the key part. Analyze the table comparing the methods on trainable parameters, training time, and performance (F1/ROC AUC). Pay close attention to the conclusions drawn about which methods perform best under which circumstances.

Summary of Findings from the Benchmark

The benchmark provides several key insights:

  • Top Performers: For the given task, rs-LoRA (a variant of LoRA) and FourierFT delivered the best performance. Standard LoRA was also a strong contender. This reinforces why LoRA is often the default choice.
  • Prompt/Prefix Tuning: These methods underperformed on the smaller model used in the benchmark, confirming that their strength is more pronounced on massive-scale models. They also increase the input sequence length, which can slow down inference.
  • Selective Methods (BitFit, LayerNorm Tuning): These are extremely lightweight but offer modest gains. They are a good choice if you are severely constrained by memory or storage, but they lack the expressive power for complex adaptations.
  • Efficiency vs. Performance: There is a clear trade-off. Methods like LoRA train more parameters than Prompt Tuning but generally achieve better results. Newer methods like FourierFT aim to get the best of both worlds: top-tier performance with even fewer parameters than LoRA.

Conclusion

You have now expanded your PEFT toolkit beyond LoRA, exploring a family of methods that work by adding trainable "prompts" or "prefixes" to a frozen model.

Key Takeaways:

  • Prompt-based vs. Weight-based PEFT: You now understand the two main philosophies. LoRA modifies model weights via low-rank updates. Prompt/Prefix Tuning adds new trainable parameters that act as learned instructions for the frozen model.
  • Prompt Tuning: Adds trainable "virtual tokens" only to the input embeddings. It's extremely parameter-efficient but most effective on very large models.
  • Prefix Tuning: A more powerful version that injects trainable prefixes into every Transformer layer, allowing for more granular control over the model's computation. It generally outperforms prompt tuning.
  • The PEFT Ecosystem: The field is rich and constantly evolving. While LoRA is a robust default, knowing about methods like Adapters, BitFit, and newer approaches like FourierFT gives you a comprehensive view of the trade-offs between parameter efficiency, performance, and implementation complexity.

Preview of the Next Lesson:

So far, we have focused on how to fine-tune models efficiently. But what are we fine-tuning them for? In many cases, we want to align a model's behavior with human preferences. A critical first step in this process is to teach a model to recognize what humans prefer. In the next lesson, we will learn how to train a reward model to capture human preferences, a foundational component of modern alignment techniques like RLHF.

Can't find a good explanation? Sign up and we'll make it for you

Sign up