Skip to main content
Create your own

Benchmarking Weight-Only vs. Weight and Activation Quantization

Introduction

In our previous lessons, we've built a strong understanding of weight-only quantization. We explored algorithms like GPTQ and AWQ, focusing on how they compress model weights to reduce VRAM usage while preserving accuracy. This corresponds to quantization schemes we can label as W4A16 or W8A16, where weights are in a low-bit format but activations remain in 16-bit floating point.

Today, we address a critical question for any AI Systems Engineer: if we can quantize weights, can we also quantize activations? And more importantly, why would we? This leads us to the concept of W8A8 quantization, where both weights and activations are represented as 8-bit integers.

Your goal in this lesson is to benchmark a weight-only quantized model against one with both weight and activation quantization to compare their performance profiles. By the end, you will understand the distinct advantages each approach offers and be able to decide which strategy is appropriate based on whether your primary goal is minimizing latency or maximizing throughput.

1. Why Weight-Only Quantization Isn't a Silver Bullet for Speed

You might intuitively expect that a model with 8-bit weights (W8A16) would be significantly faster than its 16-bit (FP16) counterpart. While it does reduce the model's memory footprint and can speed up memory-bound operations, the performance gains in compute-heavy situations are often less than expected.

To understand why, we need to think about the underlying hardware. For a matrix multiplication Y = X * W, where X are the activations and W are the weights:

  • In a W8A16 scheme, the 8-bit weights (W) must be de-quantized back to 16-bit on-the-fly to be multiplied with the 16-bit activations (X). The entire computation happens using the GPU's FP16 tensor cores. The main benefit is reduced memory bandwidth from loading the smaller weights from VRAM.
  • In a W8A8 scheme, both activations and weights are 8-bit integers. This allows the GPU to use its highly optimized INT8 tensor cores, which can perform integer matrix multiplications at a much higher rate (often 2x the throughput of FP16 cores on the same hardware).

This distinction is the key to understanding the different performance profiles. To dig deeper into this critical concept, please read the following article.

LLM Compressor is here: Faster inference with vLLM

This article from Red Hat, 'LLM Compressor is here: Faster inference with vLLM', provides a clear, systems-level explanation of why activation quantization is crucial for unlocking performance gains in compute-heavy serving scenarios.

Please read the section titled 'Enabling activation quantization in vLLM'. Focus on the explanation for why weight-only quantization can fail to deliver speed improvements in production and how activating quantization unlocks the use of faster INT8 tensor cores.

As the article highlights, weight-only quantization is beneficial for latency, but for high-throughput serving where the workload becomes compute-bound, activation quantization is what truly boosts performance.

The following graph visualizes this relationship on an RTX 4090.

RTX 4090 Performance with Different Quantization Schemes
This roofline-style graph shows the achievable performance (in Tera Operations Per Second, TOPS) for different precisions based on the arithmetic intensity of the workload. Notice how INT8xINT8 (`W8A8`) offers significantly higher peak TOPS than FP16xFP16 or INT4xFP16 (`W4A16`), but this advantage is only realized in compute-intensive regions.

2. The Challenge of Activation Quantization: SmoothQuant

If W8A8 is so powerful, why isn't it the default? The reason, as we've hinted at before, is that activations are notoriously difficult to quantize. Their values are data-dependent and often feature extreme outliers, which can wreck a simple quantization scheme.

To solve this, techniques like SmoothQuant were developed. The core insight of SmoothQuant is to make quantization easier by "smoothing" the activation outliers. It does this by migrating the quantization difficulty from the activations to the weights, which are much easier to quantize.

The following video from the creators of SmoothQuant at MIT explains this elegant idea.

SmoothQuant

Let's watch a segment from the 'SmoothQuant' video to understand the problem of activation outliers and the clever solution SmoothQuant proposes.

Please watch the following two clips: The Problem (1:12 - 2:20): This section visually demonstrates the 'outlier' problem in activations that makes them difficult to quantize. The Solution (2:20 - 3:40): This section explains the core mathematical trick of SmoothQuant: scaling activations down and weights up to make both easier to quantize, without changing the final result.

In essence, SmoothQuant uses a mathematically equivalent transformation Y = X * W = (X * s⁻¹) * (s * W). It scales down the problematic activation channels (X * s⁻¹) and scales up the corresponding weight channels (s * W). This makes the activations "smoother" and easier to fit into 8 bits, while the weights, which have a smaller dynamic range to begin with, can tolerate the scaling. This entire process happens offline, so there is no runtime overhead.

3. Performance Profile Showdown: W8A16 vs. W8A8

Now that we understand the 'why' (unlocking INT8 cores) and the 'how' (e.g., SmoothQuant), let's analyze the performance profiles of these two strategies using real-world benchmark data.

We will use a comprehensive blog post from AWS that compares multiple quantization schemes. It uses the WxAy notation we've discussed. Pay close attention to the difference between W8A16 (weight-only) and W8A8 (weight-and-activation).

Accelerating LLM inference with post-training weight and activation ...

This AWS blog post, 'Accelerating LLM inference with post-training weight and activation...', provides a treasure trove of benchmark data. We will use it to build a detailed performance profile.

First, read the subsections 'W8A8' and 'W8A16' under the 'Weights and activation: A deep dive' heading. This will solidify the definitions. Next, carefully examine the tables in the following sections for the Llama-3.1-8B-Instruct model. For each metric, compare the rows for 'Llama-3.1-8B-GPTQ-W8A16' and 'Llama-3.1-8B-GPTQ-W8A8': GPU memory utilization (Section: 'GPU memory utilization'): Note the difference in memory consumption. End-to-end latency (Section: 'End-to-end latency'): How do they compare at low concurrency (C=1) vs. high concurrency (C=128)? Inter-token latency (Section: 'Inter-token latency'): Again, compare the trend from low to high concurrency. Throughput (Section: 'Throughput'): Where does each strategy excel?

Let's synthesize your findings from the AWS blog data:

Metric W8A16 (Weight-Only) W8A8 (Weight+Activation) Analysis
GPU Memory 11.3 GB 7.8 GB W8A8 wins. Quantizing activations provides additional memory savings on top of weight quantization.
E2E Latency (C=1) 5.03s 5.47s W8A16 wins (slightly). At a single batch, the overhead of quantizing activations can be higher than the dequantization cost of weights, giving W8A16 a small edge.
E2E Latency (C=128) 40.76s 38.83s W8A8 wins. At high concurrency, the compute efficiency of INT8 tensor cores starts to dominate, leading to lower overall latency under load.
ITL (C=1) 0.020s 0.020s Tie. At the single-token generation level for a single user, the performance is virtually identical.
Throughput (C=128) 7.94 tokens/sec 8.26 tokens/sec W8A8 wins. This is the key takeaway. W8A8's superior compute efficiency translates directly to higher throughput under concurrent load.

This analysis clearly reveals the trade-off. W8A16 is optimized for low-latency, single-request scenarios, while W8A8 is designed for high-throughput, multi-user serving.

To see this visualized in another real-world scenario, examine the benchmark from the Red Hat article you read earlier.

LLM Compressor is here: Faster inference with vLLM

Let's return to the Red Hat article to see a graphical representation of this trade-off.

Please review the section 'Activation quantization performance in vLLM' and focus on Figure 3. This chart plots Time per Output Token (TPOT, the inverse of throughput) against Queries per Second (QPS). Observe the performance curves for the w4a16 (weight-only) and w8a8 (weight+activation) models.

The graph in Figure 3 perfectly illustrates our findings. The w4a16 model starts with a slight latency advantage at very low QPS, but its performance quickly degrades as load increases. The w8a8 model, however, sustains much better performance (lower time-per-output-token) as QPS goes up, showcasing its strength in compute-bound, high-throughput environments.

Conclusion

Today, you've moved beyond weight-only quantization to understand the critical role of activation quantization. By comparing the performance profiles of W8A16 and W8A8 models, you've gained a nuanced, systems-level perspective on how to optimize for different production requirements.

Key Takeaways:

  • Weight-only (WxA16) reduces memory footprint and bandwidth, making it ideal for improving latency in memory-bound, low-concurrency workloads.
  • Weight-and-activation (WxA8) unlocks the use of faster INT8 hardware cores, making it the superior choice for maximizing throughput in compute-bound, high-concurrency workloads.
  • Techniques like SmoothQuant are necessary to overcome the challenges of quantizing activations, which have a larger and more unpredictable dynamic range than weights.
  • As an AI Systems Engineer, your choice of quantization strategy is not just about a single number; it's a deliberate decision based on the specific performance characteristics (latency vs. throughput) of your target application.

Preview of the Next Lesson:

So far in this module, we've focused on compressing the model itself. However, the model weights and activations are only part of the memory story. The other massive consumer of VRAM is the KV cache. In our final lesson for this module, we will see how these optimizations can be combined. We will combine quantization with paged KV cache optimization and measure the cumulative memory savings, demonstrating how to achieve maximum efficiency and request throughput on a single GPU.

Can't find a good explanation? Sign up and we'll make it for you

Sign up