Introduction
In our previous lesson, you successfully implemented a minimal decoder-only transformer from scratch. You now have a concrete, code-level understanding of the model's fundamental components: the embedding layer, the self-attention and feed-forward sub-layers within each transformer block, and the final output head.
This lesson builds directly on that foundation. We will move from the qualitative structure to a quantitative analysis of its size. Our goal is to answer the question: "How many parameters does a given transformer model have?" You will learn to break down the model's architecture and derive formulas for the parameter count of each component.
By the end of this lesson, you will be able to achieve the learning outcome: Write a function to programmatically compute the total parameter count of a transformer from its architectural hyperparameters. This skill is the first essential step in system design, allowing you to estimate a model's memory footprint before you even load it onto a GPU.
The Core Architectural Hyperparameters
A transformer's size and shape are defined by a small set of core hyperparameters. Understanding these is key to calculating the total number of learnable weights.
Transformer Architecture Hyperparameters: Depth, Width, Heads & FFN Guide
The article 'Transformer Architecture Hyperparameters' provides an excellent overview of the six numbers that define a transformer's architecture. Let's start by reviewing them.
Read the section titled 'The Core Hyperparameters'. Focus on the definitions of V, L, d_model, n_heads, d_head, and d_ff. Connect these back to the PyTorch modules you implemented in the last lesson. For example, d_model is the embedding_dim of your nn.Embedding layer.
To summarize, the key hyperparameters are:
V(Vocabulary Size): The number of unique tokens.L(Number of Layers): The number of transformer blocks stacked together.d_model(Model Dimension): The main hidden dimension that flows through the model.n_heads(Number of Attention Heads): The number of parallel attention operations.d_head(Head Dimension): The dimension of Q, K, and V vectors within each head. Typicallyd_model / n_heads.d_ff(Feed-Forward Dimension): The intermediate dimension in the FFN. Often4 * d_model.
With these defined, we can now calculate the parameters for each part of the model.
Calculating Parameters by Component
Let's dissect the model you built and count the parameters in each nn.Module. For this exercise, we will initially ignore bias parameters for simplicity, as they contribute a negligible amount to the total count in large models.
1. Embedding Layers
-
Token Embeddings (
nn.Embedding): This is a lookup table. For each of theVtokens in the vocabulary, we store a vector of sized_model.- Parameters =
V * d_model
- Parameters =
-
Positional Encodings: In the previous lesson, we used sinusoidal positional encodings. These are fixed and not learned, so they contribute 0 parameters. Some models use learned positional embeddings, which would add
max_sequence_length * d_modelparameters.
2. The Transformer Block
Each of the L identical blocks contains an attention mechanism and a feed-forward network.
Multi-Head Attention (MHA)
In standard MHA, the input of size d_model is projected to generate queries, keys, and values for all heads at once.
- Q, K, V Projections: For each of the
n_heads, we have Q, K, V projection matrices. However, these are typically implemented as single large linear layers for efficiency.W_q: A matrix mapping the inputx(sized_model) to the query vectors for all heads (sizen_heads * d_head = d_model). Shape:[d_model, d_model]. Parameters:d_model * d_model.W_k: Similarly, a matrix of shape[d_model, d_model]. Parameters:d_model * d_model.W_v: A matrix of shape[d_model, d_model]. Parameters:d_model * d_model.
- Output Projection (
W_o): After attention is applied, the concatenated results from all heads (sized_model) are projected back through a final linear layer. Shape:[d_model, d_model]. Parameters:d_model * d_model.
Total parameters for a standard MHA block = 4 * d_model^2.
Feed-Forward Network (FFN)
The standard FFN consists of two linear layers.
- Up-projection:
nn.Linear(d_model, d_ff). Parameters:d_model * d_ff. - Down-projection:
nn.Linear(d_ff, d_model). Parameters:d_ff * d_model.
Total parameters for a standard FFN = 2 * d_model * d_ff.
This is a good point to watch a detailed, step-by-step calculation for a real-world model. The following video walks through the numbers for GPT-3.
Let us hand-calculate how GPT-3 has a total of 175B parameters | Transformers for Vision
The video 'Let us hand-calculate how GPT-3 has a total of 175B parameters' is an excellent practical exercise. It will solidify your understanding by applying these formulas to a famous, large-scale model.
Watch the following sections to see how the parameter math works for each component: Input Embeddings (6:47 - 11:53): See the calculation for both token and positional embeddings. Multi-Head Attention (12:35 - 18:35): Follow the breakdown of Q, K, V, and the output projection matrices. Feed-Forward Network (19:53 - 24:00): Observe how the two linear layers in the FFN contribute the most parameters within a block. Putting it Together (24:00 - 25:27): See how the parameters for one block are calculated and then multiplied by the number of layers. Note that the video includes bias terms in its calculation, which we will incorporate into our final function.
A Note on Modern FFNs (SwiGLU)
The standard FFN (2 * d_model * d_ff) is not the only design. Modern models like Llama use a gated variant called SwiGLU.

As the diagram shows, SwiGLU uses three matrices instead of two. To maintain a similar parameter count to the standard FFN (which typically has d_ff = 4 * d_model), SwiGLU-based models use a smaller intermediate dimension d_ff. The common choice is d_ff ≈ 2/3 * (4 * d_model) = 8/3 * d_model.
- Standard FFN Params:
2 * d_model * (4 * d_model) = 8 * d_model^2 - SwiGLU FFN Params:
3 * d_model * (8/3 * d_model) = 8 * d_model^2
The parameter counts are nearly identical, but the gated architecture often yields better performance. This is a great example of how architectural choices directly influence the parameter formulas.
A Programmatic Approach
Now, let's consolidate this logic into a Python function. Your background in computer science makes you well-suited to appreciate a programmatic and generalizable solution over simple hand-calculation.
Transformer Architecture Hyperparameters: Depth, Width, Heads & FFN Guide
The article 'Transformer Architecture Hyperparameters' culminates in a comprehensive Python function to calculate parameters. This is an excellent, production-quality example that handles many modern architectural details.
Read the section 'Total Parameter Calculation' and carefully study the full_parameter_count function. Notice how it: Breaks down the calculation by component (embedding, attention_per_layer, ffn_per_layer). Handles different ffn_type ('standard' vs 'swiglu'). Accounts for modern optimizations like Grouped-Query Attention (GQA) via the n_kv_heads argument (when n_kv_heads < n_heads). Considers tie_embeddings and optionally includes biases.
Here is the full_parameter_count function from the resource for your reference. We will use it to verify our understanding.
def full_parameter_count(
V, L, d_model, n_heads, d_ff,
n_kv_heads=None, # For Grouped-Query Attention
ffn_type="standard",
tie_embeddings=True,
include_biases=True, # Let's be precise
):
"""
Calculate total parameters for a transformer language model.
"""
if n_kv_heads is None:
n_kv_heads = n_heads
d_head = d_model // n_heads
params = {}
# 1. Embedding layer
params["embedding"] = V * d_model
# 2. Attention parameters per layer
q_params = n_heads * d_head * d_model
k_params = n_kv_heads * d_head * d_model
v_params = n_kv_heads * d_head * d_model
o_params = n_heads * d_head * d_model
params["attention_per_layer"] = q_params + k_params + v_params + o_params
if include_biases:
# Bias for Q, K, V, and O projections
params["attention_per_layer"] += (n_heads + 2*n_kv_heads) * d_head + d_model
# 3. FFN parameters per layer
if ffn_type == "standard":
ffn_weights = 2 * d_model * d_ff
ffn_biases = d_ff + d_model if include_biases else 0
else: # SwiGLU
ffn_weights = 3 * d_model * d_ff
ffn_biases = 0 # SwiGLU typically doesn't use biases
params["ffn_per_layer"] = ffn_weights + ffn_biases
# 4. Layer norms (2 per layer, each with scale and bias)
params["norm_per_layer"] = 2 * (2 * d_model) if include_biases else 2 * d_model
# Total per layer
params["total_per_layer"] = (
params["attention_per_layer"]
+ params["ffn_per_layer"]
+ params["norm_per_layer"]
)
# All layers
params["all_layers"] = L * params["total_per_layer"]
# Final layer norm + output projection
params["final_norm"] = 2 * d_model if include_biases else d_model
params["output_head"] = 0 if tie_embeddings else d_model * V
# Grand Total
params["total"] = (
params["embedding"]
+ params["all_layers"]
+ params["final_norm"]
+ params["output_head"]
)
return params
# --- Verification with GPT-2 Small (124M) ---
gpt2_small_config = {
"V": 50257, "L": 12, "d_model": 768, "n_heads": 12,
"d_ff": 3072, "ffn_type": "standard", "tie_embeddings": True
}
params = full_parameter_count(**gpt2_small_config)
print(f"Calculated GPT-2 Small parameters: {params['total'] / 1e6:.2f}M")
# Expected output: ~124.44M
The Rule of Thumb: A Simplified Formula
While the Python function gives an exact number, it's often useful to have a simplified formula for quick estimates. For large models, almost all parameters are in the transformer layers, and within each layer, the FFN and attention projections dominate.
Let's derive this simplified formula, ignoring biases and other small terms, and assuming a standard MHA and FFN (d_ff = 4 * d_model).
Parameters per layer ≈ (Attention parameters) + (FFN parameters)
- Attention ≈
4 * d_model^2 - FFN ≈
2 * d_model * d_ff=2 * d_model * (4 * d_model)=8 * d_model^2
Total parameters per layer ≈ 4 * d_model^2 + 8 * d_model^2 = 12 * d_model^2
Total model parameters ≈ L * 12 * d_model^2
This powerful approximation is presented in the "Transformer Memory Arithmetic" blog post.
This post provides a highly condensed mathematical view of a transformer's memory usage, including a simple but powerful equation for the parameter count.
Examine the equation presented in this short section. You can now see how it is derived from the component-wise breakdown we just performed: N_layers * (12 * D_model^2) is the dominant term.
This simple rule reveals a critical insight: parameter count scales linearly with the number of layers (L) but quadratically with the model's width (d_model). Doubling the depth of a model doubles its size, but doubling its width quadruples it. This is a fundamental trade-off in model design.
Conclusion
You have now successfully deconstructed a transformer's architecture to precisely calculate its parameter count. This moves you from being a user of a black-box model to an engineer who can reason about its internal structure and size from first principles.
Key Takeaways:
- A transformer's size is determined by a handful of core hyperparameters (
V,L,d_model,d_ff, etc.). - The total parameter count is the sum of the parameters from its constituent parts: embeddings, attention projections, and FFNs.
- The FFN layers are the most parameter-heavy component, typically containing twice as many parameters as the attention mechanism in each block.
- A model's size can be approximated by the formula
L * 12 * d_model^2, highlighting the quadratic scaling with model width (d_model).
Preview of the Next Lesson:
Knowing the number of parameters is the first half of the puzzle. The next question is, "How much memory do these parameters actually occupy on a GPU?" In the next lesson, "Calculate the VRAM required to store model weights for a given parameter count and data type (fp32, fp16, bf16)," we will connect parameter counts directly to VRAM usage, bringing you one step closer to answering: "Why is my 70B model using 120 GB of VRAM?"