Skip to main content
Create your own

Calculating Transformer Parameters

Introduction

In our previous lesson, you successfully implemented a minimal decoder-only transformer from scratch. You now have a concrete, code-level understanding of the model's fundamental components: the embedding layer, the self-attention and feed-forward sub-layers within each transformer block, and the final output head.

This lesson builds directly on that foundation. We will move from the qualitative structure to a quantitative analysis of its size. Our goal is to answer the question: "How many parameters does a given transformer model have?" You will learn to break down the model's architecture and derive formulas for the parameter count of each component.

By the end of this lesson, you will be able to achieve the learning outcome: Write a function to programmatically compute the total parameter count of a transformer from its architectural hyperparameters. This skill is the first essential step in system design, allowing you to estimate a model's memory footprint before you even load it onto a GPU.

The Core Architectural Hyperparameters

A transformer's size and shape are defined by a small set of core hyperparameters. Understanding these is key to calculating the total number of learnable weights.

Transformer Architecture Hyperparameters: Depth, Width, Heads & FFN Guide

The article 'Transformer Architecture Hyperparameters' provides an excellent overview of the six numbers that define a transformer's architecture. Let's start by reviewing them.

Read the section titled 'The Core Hyperparameters'. Focus on the definitions of V, L, d_model, n_heads, d_head, and d_ff. Connect these back to the PyTorch modules you implemented in the last lesson. For example, d_model is the embedding_dim of your nn.Embedding layer.

To summarize, the key hyperparameters are:

  • V (Vocabulary Size): The number of unique tokens.
  • L (Number of Layers): The number of transformer blocks stacked together.
  • d_model (Model Dimension): The main hidden dimension that flows through the model.
  • n_heads (Number of Attention Heads): The number of parallel attention operations.
  • d_head (Head Dimension): The dimension of Q, K, and V vectors within each head. Typically d_model / n_heads.
  • d_ff (Feed-Forward Dimension): The intermediate dimension in the FFN. Often 4 * d_model.

With these defined, we can now calculate the parameters for each part of the model.

Calculating Parameters by Component

Let's dissect the model you built and count the parameters in each nn.Module. For this exercise, we will initially ignore bias parameters for simplicity, as they contribute a negligible amount to the total count in large models.

1. Embedding Layers

  • Token Embeddings (nn.Embedding): This is a lookup table. For each of the V tokens in the vocabulary, we store a vector of size d_model.

    • Parameters = V * d_model
  • Positional Encodings: In the previous lesson, we used sinusoidal positional encodings. These are fixed and not learned, so they contribute 0 parameters. Some models use learned positional embeddings, which would add max_sequence_length * d_model parameters.

2. The Transformer Block

Each of the L identical blocks contains an attention mechanism and a feed-forward network.

Multi-Head Attention (MHA)

In standard MHA, the input of size d_model is projected to generate queries, keys, and values for all heads at once.

  • Q, K, V Projections: For each of the n_heads, we have Q, K, V projection matrices. However, these are typically implemented as single large linear layers for efficiency.
    • W_q: A matrix mapping the input x (size d_model) to the query vectors for all heads (size n_heads * d_head = d_model). Shape: [d_model, d_model]. Parameters: d_model * d_model.
    • W_k: Similarly, a matrix of shape [d_model, d_model]. Parameters: d_model * d_model.
    • W_v: A matrix of shape [d_model, d_model]. Parameters: d_model * d_model.
  • Output Projection (W_o): After attention is applied, the concatenated results from all heads (size d_model) are projected back through a final linear layer. Shape: [d_model, d_model]. Parameters: d_model * d_model.

Total parameters for a standard MHA block = 4 * d_model^2.

Feed-Forward Network (FFN)

The standard FFN consists of two linear layers.

  • Up-projection: nn.Linear(d_model, d_ff). Parameters: d_model * d_ff.
  • Down-projection: nn.Linear(d_ff, d_model). Parameters: d_ff * d_model.

Total parameters for a standard FFN = 2 * d_model * d_ff.

This is a good point to watch a detailed, step-by-step calculation for a real-world model. The following video walks through the numbers for GPT-3.

Let us hand-calculate how GPT-3 has a total of 175B parameters | Transformers for Vision

The video 'Let us hand-calculate how GPT-3 has a total of 175B parameters' is an excellent practical exercise. It will solidify your understanding by applying these formulas to a famous, large-scale model.

Watch the following sections to see how the parameter math works for each component: Input Embeddings (6:47 - 11:53): See the calculation for both token and positional embeddings. Multi-Head Attention (12:35 - 18:35): Follow the breakdown of Q, K, V, and the output projection matrices. Feed-Forward Network (19:53 - 24:00): Observe how the two linear layers in the FFN contribute the most parameters within a block. Putting it Together (24:00 - 25:27): See how the parameters for one block are calculated and then multiplied by the number of layers. Note that the video includes bias terms in its calculation, which we will incorporate into our final function.

A Note on Modern FFNs (SwiGLU)

The standard FFN (2 * d_model * d_ff) is not the only design. Modern models like Llama use a gated variant called SwiGLU.

Comparison of FeedForward Networks: Original Transformer vs. Llama1
A comparison between the standard FFN architecture (left) and the SwiGLU FFN used in models like Llama (right). The standard FFN uses two projection matrices. The SwiGLU variant uses three matrices: one for the gate, one for the content ('up'), and one to project down.

As the diagram shows, SwiGLU uses three matrices instead of two. To maintain a similar parameter count to the standard FFN (which typically has d_ff = 4 * d_model), SwiGLU-based models use a smaller intermediate dimension d_ff. The common choice is d_ff ≈ 2/3 * (4 * d_model) = 8/3 * d_model.

  • Standard FFN Params: 2 * d_model * (4 * d_model) = 8 * d_model^2
  • SwiGLU FFN Params: 3 * d_model * (8/3 * d_model) = 8 * d_model^2

The parameter counts are nearly identical, but the gated architecture often yields better performance. This is a great example of how architectural choices directly influence the parameter formulas.

A Programmatic Approach

Now, let's consolidate this logic into a Python function. Your background in computer science makes you well-suited to appreciate a programmatic and generalizable solution over simple hand-calculation.

Transformer Architecture Hyperparameters: Depth, Width, Heads & FFN Guide

The article 'Transformer Architecture Hyperparameters' culminates in a comprehensive Python function to calculate parameters. This is an excellent, production-quality example that handles many modern architectural details.

Read the section 'Total Parameter Calculation' and carefully study the full_parameter_count function. Notice how it: Breaks down the calculation by component (embedding, attention_per_layer, ffn_per_layer). Handles different ffn_type ('standard' vs 'swiglu'). Accounts for modern optimizations like Grouped-Query Attention (GQA) via the n_kv_heads argument (when n_kv_heads < n_heads). Considers tie_embeddings and optionally includes biases.

Here is the full_parameter_count function from the resource for your reference. We will use it to verify our understanding.

def full_parameter_count(
 V, L, d_model, n_heads, d_ff,
 n_kv_heads=None, # For Grouped-Query Attention
 ffn_type="standard",
 tie_embeddings=True,
 include_biases=True, # Let's be precise
):
    """
    Calculate total parameters for a transformer language model.
    """
    if n_kv_heads is None:
        n_kv_heads = n_heads

    d_head = d_model // n_heads
    params = {}




    # 1. Embedding layer
    params["embedding"] = V * d_model




    # 2. Attention parameters per layer
    q_params = n_heads * d_head * d_model
    k_params = n_kv_heads * d_head * d_model
    v_params = n_kv_heads * d_head * d_model
    o_params = n_heads * d_head * d_model

    params["attention_per_layer"] = q_params + k_params + v_params + o_params
    if include_biases:



        # Bias for Q, K, V, and O projections
        params["attention_per_layer"] += (n_heads + 2*n_kv_heads) * d_head + d_model




    # 3. FFN parameters per layer
    if ffn_type == "standard":
        ffn_weights = 2 * d_model * d_ff
        ffn_biases = d_ff + d_model if include_biases else 0
    else: # SwiGLU
        ffn_weights = 3 * d_model * d_ff
        ffn_biases = 0 # SwiGLU typically doesn't use biases

    params["ffn_per_layer"] = ffn_weights + ffn_biases




    # 4. Layer norms (2 per layer, each with scale and bias)
    params["norm_per_layer"] = 2 * (2 * d_model) if include_biases else 2 * d_model




    # Total per layer
    params["total_per_layer"] = (
        params["attention_per_layer"]
        + params["ffn_per_layer"]
        + params["norm_per_layer"]
    )




    # All layers
    params["all_layers"] = L * params["total_per_layer"]




    # Final layer norm + output projection
    params["final_norm"] = 2 * d_model if include_biases else d_model
    params["output_head"] = 0 if tie_embeddings else d_model * V




    # Grand Total
    params["total"] = (
        params["embedding"]
        + params["all_layers"]
        + params["final_norm"]
        + params["output_head"]
    )

    return params




# --- Verification with GPT-2 Small (124M) ---
gpt2_small_config = {
    "V": 50257, "L": 12, "d_model": 768, "n_heads": 12,
    "d_ff": 3072, "ffn_type": "standard", "tie_embeddings": True
}

params = full_parameter_count(**gpt2_small_config)
print(f"Calculated GPT-2 Small parameters: {params['total'] / 1e6:.2f}M")



# Expected output: ~124.44M

The Rule of Thumb: A Simplified Formula

While the Python function gives an exact number, it's often useful to have a simplified formula for quick estimates. For large models, almost all parameters are in the transformer layers, and within each layer, the FFN and attention projections dominate.

Let's derive this simplified formula, ignoring biases and other small terms, and assuming a standard MHA and FFN (d_ff = 4 * d_model).

Parameters per layer ≈ (Attention parameters) + (FFN parameters)

  • Attention ≈ 4 * d_model^2
  • FFN ≈ 2 * d_model * d_ff = 2 * d_model * (4 * d_model) = 8 * d_model^2

Total parameters per layer ≈ 4 * d_model^2 + 8 * d_model^2 = 12 * d_model^2

Total model parameters ≈ L * 12 * d_model^2

This powerful approximation is presented in the "Transformer Memory Arithmetic" blog post.

Transformer Memory Arithmetic

This post provides a highly condensed mathematical view of a transformer's memory usage, including a simple but powerful equation for the parameter count.

Examine the equation presented in this short section. You can now see how it is derived from the component-wise breakdown we just performed: N_layers * (12 * D_model^2) is the dominant term.

This simple rule reveals a critical insight: parameter count scales linearly with the number of layers (L) but quadratically with the model's width (d_model). Doubling the depth of a model doubles its size, but doubling its width quadruples it. This is a fundamental trade-off in model design.

Conclusion

You have now successfully deconstructed a transformer's architecture to precisely calculate its parameter count. This moves you from being a user of a black-box model to an engineer who can reason about its internal structure and size from first principles.

Key Takeaways:

  • A transformer's size is determined by a handful of core hyperparameters (V, L, d_model, d_ff, etc.).
  • The total parameter count is the sum of the parameters from its constituent parts: embeddings, attention projections, and FFNs.
  • The FFN layers are the most parameter-heavy component, typically containing twice as many parameters as the attention mechanism in each block.
  • A model's size can be approximated by the formula L * 12 * d_model^2, highlighting the quadratic scaling with model width (d_model).

Preview of the Next Lesson:

Knowing the number of parameters is the first half of the puzzle. The next question is, "How much memory do these parameters actually occupy on a GPU?" In the next lesson, "Calculate the VRAM required to store model weights for a given parameter count and data type (fp32, fp16, bf16)," we will connect parameter counts directly to VRAM usage, bringing you one step closer to answering: "Why is my 70B model using 120 GB of VRAM?"

Can't find a good explanation? Sign up and we'll make it for you

Sign up