Skip to main content
Create your own

Loading and Serving Small Models with Hugging Face Transformers

Introduction

Welcome back. In the previous lesson, you successfully configured a local, GPU-accelerated development environment using Conda and PyTorch. With that crucial foundation in place, we can now move on to the main event: running a language model.

This lesson directly addresses the learning outcome: to load and serve a small model (e.g., Phi-3-mini) using Hugging Face Transformers. We will select an appropriate small model, walk through the essential components of the Hugging Face transformers library, and write the Python code to download the model, load it onto your GPU, and generate text.

This is the "Hello, World!" of LLM serving—a fundamental skill that establishes the baseline for all the performance profiling and optimization work we'll undertake in subsequent lessons.

Selecting Our Workhorse: Phi-3-mini

For local development, profiling, and experimentation, starting with a "small" model (typically under 7 billion parameters) is far more efficient than wrestling with a 70B giant. It allows for faster iteration, requires less VRAM, and lets us focus on the systems concepts without being bottlenecked by hardware limitations.

For this course, we will use Microsoft's Phi-3-mini-4k-instruct as our primary small model. It's a 3.8 billion parameter model known for its strong performance relative to its size, making it an excellent candidate for running on consumer GPUs.

Let's start by getting familiar with the model from its official source, the Hugging Face Hub.

microsoft/Phi-3-mini-4k-instruct - Hugging Face

The model card on Hugging Face is the canonical source of information for a model. It provides details on its architecture, intended use, and, most importantly for us, how to use it.

Please review the model card. Specifically focus on: The 'Model Summary' to confirm its size (3.8B parameters). The 'How to Use' section to see the required packages like transformers and accelerate. The 'Chat Format' section, which shows the specific template the model expects for prompts. The 'Sample inference code', which we will be adapting for our own script.

The Hugging Face transformers Toolkit

As you saw in the model card, the transformers library is our primary interface for working with the model. For our purposes, three components are fundamental:

  1. AutoTokenizer: This class is responsible for converting human-readable text (strings) into a sequence of integer IDs (tokens) that the model can understand, and vice-versa. Every model has its own specific tokenizer.
  2. AutoModelForCausalLM: This is the workhorse for loading the model itself. The ForCausalLM part specifies that we're loading a decoder-only model designed for text generation (i.e., predicting the next token). It handles downloading the model's architecture and its trained weights.
  3. pipeline: A high-level helper that abstracts away much of the boilerplate code. It bundles the tokenizer, model, and the generation loop into a simple, callable object. For our goal today, the "text-generation" pipeline is exactly what we need.

Hands-On: Serving Phi-3-mini on Your GPU

Now, let's write the code to bring the model to life. The following script will load the Phi-3-mini model and tokenizer, place the model onto your GPU, and perform a generation task.

1. Environment and Dependencies

First, ensure your Conda environment from the previous lesson is active:

conda activate llm-systems

Next, install the necessary libraries for running Phi-3, as specified in the model card you just reviewed. accelerate is a key library from Hugging Face that simplifies running PyTorch models on various hardware setups, including single-GPU, multi-GPU, and TPU.

pip install transformers accelerate flash_attn --upgrade

Note: flash_attn is an optional but highly recommended dependency for optimized attention calculation, which we will cover in depth later.

2. The Inference Script

Create a new Python file named serve_phi3.py and add the following code. Each part is explained in the comments.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import time




# Ensure you have a GPU available
if not torch.cuda.is_available():
    raise SystemExit("No CUDA-enabled GPU found. Please check your setup.")

print("GPU is available. Proceeding with model loading.")




# --- 1. Define Model and Load Components ---
# The model identifier on the Hugging Face Hub.
model_id = "microsoft/Phi-3-mini-4k-instruct"




# Load the tokenizer. This converts text to tokens.
# `trust_remote_code=True` is often required for new model architectures.
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)




# Load the model itself.
# This is the most resource-intensive step.
print(f"Loading model: {model_id}")
start_time = time.time()

model = AutoModelForCausalLM.from_pretrained(
    model_id,



    # `device_map="auto"` will automatically place the model on the available GPU.
    device_map="auto",



    # `torch_dtype="auto"` lets transformers choose the optimal data type (e.g., bfloat16)
    # for your hardware, which is crucial for memory efficiency.
    torch_dtype="auto",
    trust_remote_code=True,
)

end_time = time.time()
print(f"Model loaded in {end_time - start_time:.2f} seconds.")




# --- 2. Create the Inference Pipeline ---
# The pipeline simplifies the process of getting a prediction.
# It handles tokenization, model forwarding, and detokenization.
pipe = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
)




# --- 3. Prepare the Prompt and Generation Arguments ---
# The prompt must follow the model's specified chat format.
messages = [
    {"role": "system", "content": "You are a helpful AI that explains complex concepts simply."},
    {"role": "user", "content": "Explain what happens when I run this Python script on my GPU."},
]




# Arguments to control the generation process.
generation_args = {
    "max_new_tokens": 500,       # Limit the length of the generated response.
    "return_full_text": False,   # Only return the generated part (the assistant's reply).
    "temperature": 0.7,          # Controls creativity. 0.0 is deterministic.
    "do_sample": True,           # Enable sampling for more creative responses.
}




# --- 4. Run Inference ---
print("\nGenerating response...")
start_time = time.time()

output = pipe(messages, **generation_args)

end_time = time.time()
print(f"Inference completed in {end_time - start_time:.2f} seconds.")


```grasp
{
  "type": "exercise",
  "id": "9d4a0ad3-0788-4655-85f6-9f8151f8130d"
}

--- 5. Print the Output ---

print("\nModel Output:")
print(output[0]['generated_text'])


**3. Run the Script**

Save the file and execute it from your terminal:

```bash
python serve_phi3.py

You will see output indicating the model is downloading (on the first run) and then loading. After a short wait, the script will print the model's explanation of what it's doing. Congratulations, you are now serving an LLM on your local GPU!

Connecting Code to Systems Concepts

While running a script is straightforward, it's crucial to understand the underlying system implications. This is the core of our journey from AI Engineer to AI Systems Engineer.

VRAM: The Cost of Loading

The most significant step in our script is AutoModelForCausalLM.from_pretrained(...). This call allocates a substantial amount of GPU memory (VRAM) to hold the model's weights. How much? We can estimate it.

The following guide explains the relationship between model size, precision, and memory.

Optimizing LLMs for Speed and Memory

The article 'Optimizing LLMs for Speed and Memory' from Hugging Face provides a clear, practical guide to the memory arithmetic of LLMs.

Read the first main section, '1. Lower Precision', up to the part where it begins discussing quantization ('Now what if your GPU does not have 32 GB of VRAM?'). Focus on the rule of thumb provided for calculating memory requirements based on parameter count and precision (bfloat16/float16).

From the reading, the key rule of thumb is:

Loading the weights of a model having X billion parameters requires roughly 2 * X GB of VRAM in bfloat16/float16 precision.

Our Phi-3-mini model has 3.8B parameters. By setting torch_dtype="auto", we allow transformers to load it in bfloat16 on compatible GPUs. Therefore, we can estimate the memory required for the weights:

This calculation only accounts for the model weights. Additional memory is needed for the activations, KV cache, and the CUDA context itself. This simple calculation begins to answer the question, "Why is my model using so much VRAM?" We will make these calculations exact in the next module.

Conclusion

In this lesson, you have taken a critical practical step. You've gone from an empty, configured environment to loading a sophisticated, pre-trained language model onto your GPU and using it to generate text.

Key Takeaways:

  • Hugging Face Transformers is the standard library for loading and running open-source models, using AutoTokenizer, AutoModelForCausalLM, and pipeline as key components.
  • Loading a model is as simple as specifying its ID, but crucial arguments like device_map="auto" and torch_dtype="auto" are essential for correct GPU placement and memory-efficient loading.
  • Every model has a specific chat/prompt template that must be followed for optimal results.
  • The VRAM required to load a model is primarily determined by its parameter count and the numerical precision (e.g., bfloat16) of its weights.

Preview of the Next Lesson:

Our estimation of 7.6 GB is just a back-of-the-envelope calculation. How much memory is actually being used? Is the GPU working hard or is it idle most of the time? To answer these questions, we need to measure. In the next lesson, we will learn to profile GPU memory usage and utilization during inference using nvidia-smi and PyTorch CUDA utilities.

Can't find a good explanation? Sign up and we'll make it for you

Sign up