Skip to main content
Create your own

Model Quantization for Size & VRAM Reduction

Hello! Welcome to the first lesson in our module on Efficient AI: Deployment and Optimization.

In the previous module, we explored the cutting edge of generative AI and wrestled with the profound ethical and legal questions these powerful technologies raise. We've seen what's possible, from creating custom videos to fine-tuning models for specialized text generation. Now, we pivot from the "what" and "why" back to the "how"—specifically, how do we run these enormous models without access to a datacenter?

This lesson directly addresses the learning outcome: Apply model quantization techniques (INT8, 4-bit) to reduce model size and VRAM usage. You've likely seen terms like "4-bit model" or "GGUF" on platforms like Hugging Face. Today, you'll learn exactly what they mean and how to use them. We will cover:

  • The fundamental reason why quantization is necessary.
  • The core mathematical concepts of mapping high-precision floating-point numbers to low-precision integers.
  • A practical, code-based application of quantization using popular libraries.
  • An overview of the different methods and formats used in the field today.

Let's begin by understanding the problem that quantization solves.


1. The Problem: The Memory Cost of Precision

Large Language Models (LLMs) are defined by their parameter count, which can be in the billions. In a typical neural network, each of these parameters (weights) is stored as a 32-bit floating-point number (FP32).

A single FP32 value requires 4 bytes of memory. Let's do some quick math for an 8-billion parameter model like Llama 3 8B:

Loading just the model weights requires 32 GB of VRAM, exceeding the capacity of most consumer and even many professional-grade GPUs. This doesn't even account for the memory needed for activations during inference. This is the fundamental barrier that quantization aims to break down.

To get a practical overview of this problem and the solution, let's watch the first part of a helpful video.

Quantizing LLMs - How & Why (8-Bit, 4-Bit, GGUF & More)

To start, let's get a practical overview of why quantization is so important and what it achieves. The video 'Quantizing LLMs - How & Why' by Adam Lucek provides an excellent introduction.

Please watch the first two parts: Introduction (00:00 - 02:20): This section shows a real-world example of a large model and its quantized versions, highlighting the dramatic reduction in size and explaining the goal of running models on consumer hardware. How LLM Weights are Stored (02:20 - 12:02): This is a crucial section that reviews how numbers are represented in binary and explains the floating-point formats (FP32, FP16, BF16) that are standard for training and storing neural networks. Given your CS background, much of the binary/decimal conversion will be familiar, but pay close attention to the structure of floating-point numbers (sign, exponent, mantissa) as this is key to understanding quantization.

As the video explains, quantization reduces the memory footprint by representing the model's weights with lower-precision data types. Instead of using 32 bits for every number, we can use 16 (FP16, BF16), 8 (INT8), or even 4 (INT4). This drastically reduces the model's size and VRAM requirements, making it possible to run massive models on accessible hardware.


2. The Mathematics of Quantization

The core idea of quantization is to map a range of high-precision floating-point values to a smaller range of low-precision integer values. But how is this mapping done without losing all the information? The process relies on two key parameters: a scale factor and a zero-point.

Let's explore the two primary linear quantization strategies.

A Visual Guide to Quantization

Now that we understand the 'why', let's dive into the 'how'. The blog post 'A Visual Guide to Quantization' by Maarten Grootendorst offers exceptionally clear explanations and diagrams for this fundamental mapping process.

Please read Part 2 of the guide, focusing on the section 'Symmetric Quantization' and 'Asymmetric Quantization'. Pay close attention to: The formulas for calculating the scale factor (s) and zero-point (z). The visual difference between how the two methods map the floating-point range to the integer range. The concept of quantization error – the inevitable loss of precision when de-quantizing.

Let's summarize the two methods:

A. Symmetric Quantization (e.g., absmax)

  • Concept: Maps the floating-point range [-abs_max, +abs_max] to the integer range [-127, 127] (for INT8). The floating-point zero 0.0 maps directly to the integer 0.
  • Scale Factor (S): A single value that defines the mapping. It's calculated by dividing the absolute maximum value of the float tensor by the maximum value of the integer range. where b is the number of bits (e.g., 8).
  • Quantization:
  • Dequantization:

B. Asymmetric Quantization (e.g., zero-point)

  • Concept: Maps the full floating-point range [min, max] to the full integer range [-128, 127] (for INT8). The floating-point zero 0.0 may not map to the integer 0.
  • Scale Factor (S):
  • Zero-Point (Z): An integer offset that represents the floating-point value 0.0. It ensures that 0.0 can be perfectly represented.
  • Quantization:
  • Dequantization:

The inevitable difference between the original float value r and the dequantized value r' is the quantization error. The goal of advanced quantization methods is to minimize this error.

The Outlier Problem

A major challenge in quantization is the presence of outliers—a few weights with very large magnitudes. In absmax quantization, a single large outlier can dramatically increase the scale factor, forcing all the smaller, non-outlier values to be mapped to a very narrow range of integers. This crushes their precision and harms model performance.

More advanced techniques were developed to handle this. One famous example is the LLM.int8() method, which treats outliers differently.

Schematic of LLM.int8() Quantization
This diagram illustrates the `LLM.int8()` method. It handles the 'outlier problem' by separating large-magnitude values and processing them in 16-bit precision, while quantizing the rest of the values to 8-bit integers. This hybrid approach significantly reduces memory usage while preserving the precision of the most influential weights, thus maintaining model performance.
Test your understanding!

Imagine you have the following tensor of weights: [-3.5, -1.0, 0.0, 0.5, 2.0, 7.8]. You want to apply symmetric quantization to convert it to 4-bit integers (INT4), which have a range of [-8, 7].

  1. What is the absmax value of the tensor?
  2. What is the scale factor S?
  3. What is the quantized tensor? (Remember to round to the nearest integer).
Show answer
  1. Absmax value: The largest absolute value in the tensor is |7.8|, so α = 7.8.

  2. Scale factor (S): The float range is [-7.8, 7.8] and the integer range is [-8, 7]. We use the maximum absolute value of the integer range, which is 7.

  3. Quantized tensor: We divide each weight by S and round.

    • -3.5 / 1.114 ≈ -3.14 → -3
    • -1.0 / 1.114 ≈ -0.89 → -1
    • 0.0 / 1.114 = 0 → 0
    • 0.5 / 1.114 ≈ 0.45 → 0
    • 2.0 / 1.114 ≈ 1.79 → 2
    • 7.8 / 1.114 ≈ 7.0 → 7

    The quantized tensor is [-3, -1, 0, 0, 2, 7]. Notice the two smallest positive values (0.0 and 0.5) both got mapped to 0, demonstrating quantization error.


3. Applying Quantization with bitsandbytes

Theory is essential, but the real power comes from applying it. Fortunately, the Hugging Face ecosystem makes this incredibly straightforward, primarily through the bitsandbytes library.

Let's see how to load a model in 4-bit precision with just a few lines of code.

Quantizing LLMs - How & Why (8-Bit, 4-Bit, GGUF & More)

Theory is great, but let's see how easy it is to apply these techniques in code. We'll return to Adam Lucek's video, which demonstrates 8-bit and 4-bit quantization using the popular bitsandbytes library within Hugging Face.

Please watch these two sections: Implementation with bitsandbytes (14:59 - 18:26): Focus on the BitsAndBytesConfig and how it's passed to the from_pretrained method. Notice the dramatic reduction in memory footprint and VRAM usage he demonstrates. Performance Evaluation (18:26 - 21:51): Observe how he compares the text output of the base, 8-bit, and 4-bit models. Pay attention to the concept of perplexity as a quantitative measure of the performance degradation.

As the video shows, you can quantize a model on-the-fly during loading. Here is a typical code snippet to achieve 4-bit quantization:

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

# Define the 4-bit quantization configuration
quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",  # NormalFloat-4, a sophisticated 4-bit data type
    bnb_4bit_use_double_quant=True, # A trick to save even more memory for quantization metadata
    bnb_4bit_compute_dtype=torch.bfloat16 # Perform computations in 16-bit for stability
)

# Load the model with the specified quantization configuration
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Meta-Llama-3-8B",
    quantization_config=quantization_config,
    device_map="auto" # Automatically place the model on available hardware (e.g., GPU)
)

tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B")

The key takeaway is that by simply creating a BitsAndBytesConfig object and passing it to the from_pretrained method, the library handles all the complex mathematics of scaling and mapping for you. The result is a model that uses a fraction of the VRAM but, as the video demonstrates, maintains surprisingly good performance.


4. A Survey of the Quantization Landscape

While bitsandbytes is a common entry point, the world of quantization is vast. It's useful to be aware of the different philosophies and formats you'll encounter.

The Complete Guide to LLM Quantization - LocalLLM.in

Beyond the basic bitsandbytes implementation, there's a whole ecosystem of quantization methods and formats. 'The Complete Guide to LLM Quantization' from LocalLLM.in provides an excellent survey.

Please read the following sections to get a broader view of the landscape: Major LLM Quantization Methods: Read the summaries for GPTQ, AWQ, and especially GGUF, which is optimized for CPU usage. Types of Quantization Approaches: PTQ vs QAT: This section formally defines the difference between quantizing after training (which is what we've focused on) and integrating quantization into the training process. Understanding Quantization Naming Conventions in GGUF: This is very practical. It decodes names like Q4_K_M and explains what the letters and numbers mean.

Let's break down some of that terminology.

Post-Training Quantization (PTQ) vs. Quantization-Aware Training (QAT)

Aspect Post-Training Quantization (PTQ) Quantization-Aware Training (QAT)
When? After the model is fully trained. During the training or fine-tuning process.
Process Take a pre-trained model and convert its weights. Simple & fast. Simulate the effects of quantization during training loops.
Accuracy Good, but can suffer significant degradation at very low bits (≤4). Excellent, as the model learns to be robust to quantization noise.
Use Case The most common approach. Great for quickly deploying models. For custom models where preserving maximum accuracy is critical.
Our Example Loading with bitsandbytes is a form of PTQ. Requires a full training setup.

GGUF: The CPU Powerhouse

You've likely seen GGUF files on Hugging Face. GGUF is a file format developed by the llama.cpp project, designed for one primary purpose: to run LLMs efficiently on CPUs.

  • It's a single, portable file that contains the model architecture and all its quantized weights.
  • It supports a huge variety of quantization methods, allowing users to choose the perfect balance of size, speed, and quality for their specific hardware.

The GGUF naming convention, like Q4_K_M, tells you exactly how the model was quantized:

  • Q{Bits}: The number of bits per weight (e.g., Q4 is 4-bit, Q8 is 8-bit).
  • {Method}: The quantization technique used. K signifies that K-means clustering was used to group weights, which is a more intelligent way to create the quantization "bins" and reduces error.
  • {Size}: The block size used for quantization. S, M, and L stand for Small, Medium, and Large.

So, Q4_K_M is a popular and well-balanced choice: 4-bit precision, using the K-means method, on Medium-sized blocks of weights.


Conclusion

In this lesson, we demystified the crucial technique of model quantization. We moved from the high-level problem of massive model sizes to the low-level mathematics of number representation and finally to the practical application of loading and running a 4-bit model.

Key Takeaways:

  • Quantization is essential for running large models on consumer hardware, reducing model size and VRAM usage by using low-precision data types (like INT8/INT4) instead of FP32.
  • The core of quantization is a mathematical mapping from a float range to an integer range, defined by a scale factor and often a zero-point.
  • Symmetric and Asymmetric are two fundamental linear quantization approaches, each with its own way of defining this mapping.
  • Libraries like bitsandbytes make Post-Training Quantization (PTQ) incredibly accessible, allowing on-the-fly conversion of Hugging Face models.
  • The GGUF format, used by llama.cpp, is a highly optimized standard for running quantized models on CPUs, with a descriptive naming system that indicates the exact quantization method used.

Preview of the Next Lesson:

Quantization is a way to shrink a model by compressing its existing weights. But what if we could train a smaller model from the start to be just as smart as a big one? In our next lesson, we will explore knowledge distillation, a technique where a small "student" model learns to mimic the behavior of a large "teacher" model, offering another powerful path toward efficient AI.

Can't find a good explanation? Sign up and we'll make it for you

Sign up