Introduction
In our last lesson, we demystified the process of counting a transformer's parameters, culminating in a Python function and a powerful rule of thumb (L * 12 * d_model^2) to estimate a model's size. You can now look at a model's hyperparameters and determine how many millions or billions of parameters it contains.
This lesson takes the crucial next step. We will translate that abstract parameter count into a concrete, physical resource requirement: GPU VRAM. This directly addresses one of the core questions of this course: "Why is my 70B model using 120 GB of VRAM?"
Our goal is to master the learning outcome: Calculate the VRAM required to store model weights for a given parameter count and data type (fp32, fp16, bf16). This calculation forms the absolute baseline for any LLM memory analysis.
The Language of Precision: From Parameters to Bytes
The number of parameters alone doesn't tell us the memory footprint. The missing piece is the numerical precision used to store each parameter. Each parameter in a model is a floating-point number, and the amount of memory it occupies depends on its data type.
The most common data types in deep learning are:
- fp32 (single precision): The standard 32-bit floating-point format.
- fp16 (half precision): A 16-bit format that saves memory at the cost of reduced precision and range.
- bf16 (brain float): Another 16-bit format, developed by Google, which prioritizes a large dynamic range (like fp32) over precision.
The image below shows how the 32 or 16 bits are allocated for each type.

To fully grasp the implications of these differences, let's dive into their structure. Given your CS background, a bit-level understanding will provide a solid foundation for why these types behave differently, especially when we discuss quantization later.
What are Float32, Float16 and BFloat16 Data Types?
The video 'What are Float32, Float16 and BFloat16 Data Types?' offers a clear explanation of how these formats are constructed and the trade-offs involved.
Watch the following segments: Float32 Structure (0:54 - 3:23): Understand the standard 32-bit layout (sign, exponent, mantissa). Float16 and BFloat16 Structures (3:23 - 4:46): Pay close attention to how fp16 and bf16 allocate their 16 bits differently. Range Comparison (4:46 - 7:35): This is the key part. Notice why converting between fp32 and bf16 is simpler due to their similar exponent range, which is a major reason for bf16's adoption in training and inference.
The crucial takeaway for our VRAM calculation is the number of bytes each parameter occupies:
- fp32: 32 bits = 4 bytes
- fp16: 16 bits = 2 bytes
- bf16: 16 bits = 2 bytes
The Fundamental Formula for Model Weight Memory
With the bytes-per-parameter established, the calculation for storing the model weights becomes straightforward. You simply multiply the number of parameters by the bytes required for the chosen data type.
What is GPU Memory and Why it Matters for LLM Inference
The article 'What is GPU Memory and Why it Matters for LLM Inference' by BentoML clearly outlines the components of VRAM usage. Let's focus on the first and most fundamental component: model weights.
Read the sections 'Model weights' and the 'Note' immediately following it. The text provides the direct formula and lists the bytes per parameter for various data types, which we will now put into practice.
As the article states, the formula is:
To convert this to gigabytes (GB), you divide by . However, a common and useful convention in the field is to approximate 1 billion bytes as 1 GB. We will use this simpler convention for our estimates.
Calculation in Practice: The 70B Model Case
Let's apply this to answer one of the course's framing questions. A 70 billion parameter model is a standard large model size (e.g., Llama-2 70B).
Scenario 1: Using Full Precision (fp32)
- Parameters: 70,000,000,000
- Bytes per parameter: 4
- Calculation:
70B params * 4 bytes/param = 280B bytes - VRAM Required: ~280 GB
Scenario 2: Using Half Precision (fp16 or bf16)
- Parameters: 70,000,000,000
- Bytes per parameter: 2
- Calculation:
70B params * 2 bytes/param = 140B bytes - VRAM Required: ~140 GB
This simple calculation reveals a critical insight: even with half-precision, a 70B model's weights alone require 140 GB of VRAM. This is why a single NVIDIA A100 GPU with 80 GB of VRAM cannot load the model without further optimizations like quantization or model parallelism, which we will cover in future modules.
Overhead and Practical Estimation Tools
The VRAM needed for model weights is the non-negotiable baseline. However, it's not the only consumer of memory. During inference, other components also require VRAM:
- Activations: Intermediate results from calculations within the model layers.
- KV Cache: Stores attention states for previously processed tokens.
- Framework Overhead: Memory used by PyTorch, CUDA kernels, etc.
For a quick but more holistic estimate, a common practice is to add a small overhead margin (e.g., 20%) to the weight memory.
GPU VRAM Calculation for LLM Inference and Training
The video 'GPU VRAM Calculation for LLM Inference and Training' demonstrates this practical approach and introduces tools that can automate these estimations.
Watch the following segments: Formula with Overhead (6:39 - 8:20): Observe how the presenter incorporates a 1.2x multiplier to account for overhead and applies it to a 7B model. VRAM Estimator Tool (8:54 - 11:57): See how an online tool can be used to quickly estimate VRAM. Notice that the largest portion of the memory reported by the tool is for 'parameters', which corresponds to the calculation we've just learned.
While we will deconstruct each overhead component (activations, KV cache) in subsequent lessons, this gives you a complete first-pass mental model for VRAM estimation. For this lesson, our focus remains on precisely calculating the weight component.
Self-Check Exercise
Let's test your understanding. Calculate the VRAM required to store the weights of a 13B parameter model (like Llama-2 13B) using the following data types. We'll also include INT8 as a preview for our upcoming lessons on quantization.
fp32(4 bytes)fp16(2 bytes)INT8(1 byte)
...
Answers:
- fp32:
13B * 4 bytes = 52 GB - fp16:
13B * 2 bytes = 26 GB - INT8:
13B * 1 byte = 13 GB
Notice how moving from fp16 to INT8 halves the memory requirement again. This is the power of quantization. For quick reference, the table below (from the Propelrc blog LINK) shows typical VRAM requirements for various model sizes, which aligns with our calculations.
| Model Size | FP16 VRAM | INT8 VRAM | INT4 VRAM |
|---|---|---|---|
| 3B parameters | 6-8GB | 3-4GB | 1.5-2GB |
| 7B parameters | 14-16GB | 7-8GB | 3.5-4GB |
| 13B parameters | 26-30GB | 13-15GB | 6.5-7.5GB |
| 70B parameters | 140-150GB | 70-75GB | 35-38GB |
Conclusion
In this lesson, you've connected the abstract concept of parameters to the physical reality of VRAM consumption. You can now perform the most fundamental calculation in LLM systems engineering.
Key Takeaways:
- VRAM for model weights is calculated as:
Number of Parameters * Bytes per Parameter. - The choice of data type is critical:
fp32uses 4 bytes, whilefp16andbf16use 2 bytes, instantly halving the memory footprint. - This weight memory is the static baseline. Total VRAM usage will be higher due to dynamic components like activations and the KV cache.
- Even for moderately sized models (13B+), fp16 precision is essential to fit on most single consumer or enterprise GPUs.
Preview of the Next Lesson:
We've now quantified the memory that is statically occupied by just loading the model. But what happens when we perform a forward pass? The moment we feed data into the model, a new, dynamic memory consumer appears: activations. In our next lesson, we will learn to calculate the activation memory consumed during a forward pass for a given batch size and sequence length, adding another critical piece to our VRAM puzzle.