Skip to main content
Create your own

Supervised Fine-Tuning with Custom Instructions

Hello! Welcome to a new module focused on the cutting edge of AI: Large Language Models.

In our previous module, we concluded our journey through deep reinforcement learning by implementing Proximal Policy Optimization (PPO). We saw how its clever "clipping" mechanism provides the stability needed to train robust agents, making it one of the most successful algorithms in the field. This foundation in policy-based learning will be surprisingly relevant as we move forward.

Today, we begin our exploration of how to shape the behavior of massive, pre-trained language models. A base LLM is like a genius who has read the entire internet but has no specific job training. It knows facts, but it doesn't know how to be a helpful assistant. Our first step in training it for a specific role is Supervised Fine-Tuning.

Your learning outcome for this lesson is to perform supervised fine-tuning (SFT) on a custom instruction dataset. We will cover the entire workflow, from creating a high-quality dataset from scratch to using modern tools to efficiently train a model.

1. Positioning SFT in the LLM Lifecycle

Before we dive in, let's understand where SFT fits into the bigger picture. The journey of creating a powerful, aligned LLM like ChatGPT or Llama 3 typically involves several stages.

LLM Lifecycle: Pre-Training, Fine-Tuning, and In-Context Learning
This diagram shows the complete lifecycle of a Large Language Model. We are focusing on the **Supervised Fine-Tuning (SFT)** stage, which takes a general pre-trained model and teaches it to follow specific instructions, transforming it into a helpful assistant.

As the diagram illustrates, SFT is the bridge between a general-purpose pre-trained model and a specialized, instruction-following model. It's the process of teaching the model a specific behavior or skill.

Supervised Fine-Tuning with SmolLM3

Let's start with a clear definition of Supervised Fine-Tuning and why it's so effective. This article from Hugging Face provides an excellent conceptual foundation.

Please read the first three sections: 'What is Supervised Fine-Tuning?', 'Why SFT Works: The Science Behind It', and the short paragraph under 'The SmolLM3 SFT Journey'. Focus on understanding that SFT is not about teaching the model new factual knowledge, but rather about 'reshaping how existing knowledge is applied' and adapting its behavior.

To put it simply, SFT works by showing the model examples of high-quality "prompt" and "completion" pairs. The model then learns to generate completions that match the style, format, and intent of the examples it has been shown.

Supervised Fine-tuning Process Overview
This image contrasts pre-training with SFT. Pre-training uses vast amounts of raw text to build general language understanding. SFT, on the other hand, uses curated, high-quality demonstration data (prompt-completion pairs) to teach a pre-trained base model how to follow instructions.

2. The Core Ingredient: A Custom Instruction Dataset

The learning outcome specifically mentions a custom instruction dataset. This is, without a doubt, the most critical element for successful fine-tuning. The quality of your dataset directly determines the quality of your fine-tuned model. Garbage in, garbage out.

But when and why would you need to build your own dataset instead of using one of the many publicly available ones? And how would you go about it?

The following resource provides a complete, real-world walkthrough of creating a custom instruction dataset for a specific domain—in this case, for a tool called Firecrawl. This is a perfect example of the process you'd follow in a professional setting.

How to Create Custom Instruction Datasets for LLM Fine-tuning

This blog post from Firecrawl is an end-to-end guide on creating a custom instruction dataset. It covers the 'why', the 'what', and the 'how' with practical code examples. We will break it down into a few steps.

First, read the sections 'What Are Instruction Datasets?', 'When to Create a Custom Instruction Dataset?', and 'What Formats Do Instruction Datasets Use?'. This will establish the motivation and the standard data structures we'll be working with.

A Practical Workflow for Dataset Creation

Now that you understand the motivation, let's look at the actual process. The Firecrawl article details a four-script pipeline. We'll examine the purpose of each step to understand the end-to-end workflow.

How to Create Custom Instruction Datasets for LLM Fine-tuning

Let's continue with the same article to see how the theory is put into practice. This section provides an overview of the code-based workflow.

Skim through the sections from 'Project overview' to 'Uploading generated pairs to HuggingFace Datasets'. You don't need to read every line of code, but focus on understanding the purpose of each stage: Curating a raw dataset (scrape_raw_data.py): How do they gather the initial, domain-specific text (e.g., from documentation and blogs)? Cleaning the documents (process_dataset.py): What steps are taken to clean, chunk, and filter the raw text into usable pieces of information? Generating instruction-answer pairs (generate.py): This is a key modern technique. How do they use another powerful LLM (like GPT-4) to create synthetic instruction-answer pairs from the cleaned text chunks? Uploading to HuggingFace (upload_to_hf.py): How is the final dataset prepared and shared for easy use in training pipelines?

This process—scrape, clean, chunk, generate, format—is a powerful and scalable blueprint for creating high-quality, specialized instruction datasets for virtually any domain.

3. The SFT Training Mechanism

We have our dataset of (instruction, answer) pairs. How does the model actually learn from it?

At its core, SFT is simply next-token prediction, the same objective used during pre-training. We concatenate the instruction and the answer, and we train the model to predict the next token at each position.

However, there's a crucial detail: we only care about the model's ability to generate the answer. We don't need it to "learn" how to generate the prompt, as that will be our input during inference. Therefore, we mask the loss for the prompt tokens. The model's predictions for the prompt part of the sequence are ignored when calculating the loss and updating the weights.

The following video explains this mechanism beautifully.

Finetune LLMs to teach them ANYTHING with Huggingface and Pytorch | Step-by-step tutorial

This video from 'Neural Breakdown with AVB' provides a phenomenal 'from-scratch' explanation of the SFT loss calculation. Understanding this is key to grasping what SFT is actually doing under the hood.

Please watch from 15:48 to 30:37. This is the most important theoretical part of the lesson. 15:48 - 25:30: Pay close attention to how an input sequence and a target sequence are created by shifting the tokenized text. This is the foundation of next-token prediction. 25:30 - 28:29: This part is critical. It shows how the target labels corresponding to the initial prompt are replaced with a special value (-100) to mask them from the loss calculation. This ensures the model only learns to generate the desired completion. 28:29 - 30:37: The video demonstrates a complete training loop, showing the loss decreasing as the model learns to generate the target phrase. This connects the theory directly to a practical result.

4. SFT in Practice with Hugging Face TRL

Manually implementing the training loop, loss masking, and data handling as shown in the video is a great way to learn. However, for practical applications, the community has developed powerful tools that abstract away this complexity. The most popular is the TRL (Transformer Reinforcement Learning) library from Hugging Face, which contains the SFTTrainer.

The SFTTrainer handles all the details we just learned about automatically:

  • Applying chat templates to format prompts and completions.
  • Tokenizing the data.
  • Masking the loss on the prompt tokens.
  • Managing the training loop.

It also integrates with other crucial libraries like PEFT (Parameter-Efficient Fine-Tuning) and BitsAndBytes to make it possible to fine-tune enormous models on a single GPU.

Let's walk through a full example.

Creating your own ChatGPT: Supervised fine-tuning (SFT)

This video by Niels Rogge provides a complete, production-ready workflow for SFT using the SFTTrainer. It ties together everything we've discussed.

Watch the following segments, focusing on how the SFTTrainer and its configuration simplify the process: Data Loading & Formatting (13:48 - 28:12): See how a dataset is loaded and how a chat template is used to format the conversational data into a single string. This is the automated version of what we saw in the previous video. Efficient Fine-Tuning with QLoRA (28:12 - 36:50): Full fine-tuning is memory-intensive. This section introduces QLoRA (Quantized LoRA) as a highly efficient alternative. Understand the core idea: you load the base model in a low-precision format (4-bit) and train small 'adapter' layers on top. This is a crucial technique for practical SFT. Training with SFTTrainer (36:50 - 46:19): This is the main implementation. Observe how the SFTTrainer is instantiated with the model, dataset, tokenizer, and training configurations (SFTConfig). Notice how simple it is to start training with trainer.train(). Saving and Inference (46:19 - 51:18): Finally, see how the trained adapters are saved and how the final model is used for inference.

Test your understanding!

You have a custom dataset of 1,000 instruction-answer pairs for a specialized legal domain. A colleague suggests that during SFT, you should try to maximize the model's accuracy on predicting the tokens in both the instruction and the answer to ensure it fully understands the legal language.

Based on what you've learned, what is the flaw in this suggestion?

Show answer

The suggestion is flawed because the goal of SFT is to teach the model how to generate the answer given the instruction, not to learn how to generate the instruction itself. During inference, the instruction (prompt) is provided as input.

By calculating loss on the instruction tokens, you would be spending valuable training computation on a task that is not relevant to the model's final use case. The standard and correct practice is to mask the loss on the instruction tokens and only train the model to predict the answer tokens.

Conclusion

You have now covered the entire process of Supervised Fine-Tuning. This is the foundational technique for specializing a pre-trained LLM for any conversational or instruction-following task.

Key Takeaways:

  • SFT Teaches Behavior: SFT adapts a pre-trained model's behavior to follow instructions and generate responses in a specific format, rather than teaching it new factual knowledge.
  • Data is King: The success of SFT is almost entirely dependent on the quality, diversity, and relevance of the custom instruction dataset.
  • The SFT Workflow: A typical process involves scraping raw data, cleaning and chunking it, using an LLM to generate synthetic instruction-answer pairs, and formatting it for training.
  • Loss Masking is Crucial: The core training mechanism is next-token prediction, but the loss is only calculated on the completion/answer tokens to focus the model on its generation task.
  • Modern Tooling is Essential: Libraries like Hugging Face's TRL (with SFTTrainer) and techniques like QLoRA are indispensable for making SFT practical and efficient on available hardware.

Preview of the Next Lesson:

In this lesson, we briefly encountered QLoRA as a method for "parameter-efficient" fine-tuning. This technique is revolutionary because it allows us to fine-tune massive models without updating all of their billions of parameters. In our next lesson, we will dive deep into its core component: LoRA (Low-Rank Adaptation). We'll explore the linear algebra behind it and implement it to understand how it achieves such impressive efficiency.

Can't find a good explanation? Sign up and we'll make it for you

Sign up