Skip to main content
Create your own
Lesson illustration

Estimating System Capacity

Welcome to the next lesson in our journey through the "Foundations of Scalable Systems." In our last session, we explored the fundamental choice between vertical and horizontal scaling, concluding that sound scaling decisions aren't based on dogma but on the specific needs of a system. You can't justify moving from a simple vertical setup to a complex horizontal one without first answering the question, "How much capacity do we actually need?"

Today, we'll learn how to answer that question. We will delve into one of the most practical skills in a senior engineer's toolkit: back-of-the-envelope calculations. This is the art of using quick, approximate estimations to gauge a system's capacity requirements. Mastering this will not only be crucial for your system design interviews but will also give you the ability to quickly assess the feasibility of an architecture before writing a single line of code.

Our goal is to learn how to perform these calculations to estimate three critical metrics: Queries Per Second (QPS), storage, and bandwidth.

1. What is Back-of-the-Envelope Estimation?

Before diving into formulas, it's essential to understand the philosophy behind this technique. Back-of-the-envelope (BOTE) estimation is not about finding the exact, correct number. It's about getting an answer that is in the right order of magnitude. Are we talking about needing 10 servers or 1000? Terabytes of data or petabytes? This is the level of precision we're aiming for.

This skill is a cornerstone of the initial "Clarify Requirements" phase of a system design interview, as shown in the cheat sheet below.

Capacity estimation is a key part of Step 1 in a structured system design interview, helping to define the scale of the problem you need to solve.

Let's watch a short video from ByteByteGo that introduces the concept and its importance.

Back-Of-The-Envelope Estimation / Capacity Planning

This video provides a great overview of what BOTE math is, why it's used, and the level of accuracy expected. It also introduces some practical techniques we'll use later.

Watch the first minute to understand the core purpose of BOTE, as explained from the introduction. Then, jump ahead to watch the segment on simplifying calculations, which introduces the powerful technique of using scientific notation.

The key takeaway is that BOTE helps us make quick, high-level architectural decisions. If your math suggests you need to handle 1 million requests per second, you know from the start that a single server won't work, justifying the need for horizontal scaling.

2. The Estimation Toolkit: Foundational Numbers

To perform these calculations quickly, you need a mental toolkit of common numbers and approximations. You are not expected to have these memorized perfectly, but you should be familiar with their orders of magnitude.

Powers of 10 and 2

When dealing with data, we measure in bytes, kilobytes, megabytes, etc. In formal computer science, these are powers of 2 (e.g., 1 KB = 1024 bytes). For BOTE, we aggressively approximate them to powers of 10 to make the math easier.

Power (of 10) Approximate Value Full Name Short Name
1 Thousand 1 Kilobyte 1 KB
1 Million 1 Megabyte 1 MB
1 Billion 1 Gigabyte 1 GB
1 Trillion 1 Terabyte 1 TB
1 Quadrillion 1 Petabyte 1 PB

This approximation is a standard and expected practice in system design interviews.

Common Assumptions

You also need to make reasonable assumptions about the size of the data you're handling. Here are some good starting points.

Back To Envelope Or Capacity Estimation - Wisdom - GitBook

This guide provides an excellent list of baseline numbers that you can use to ground your estimations.

Review the table of Typical Latency to appreciate the vast performance differences between operations. You don't need to memorize it, but internalize that disk and network operations are orders of magnitude slower than memory. Then, focus on the list under Common Data Size Assumptions. These are the kinds of numbers you'll state in an interview (e.g., "I'll assume the average image size is 200 KB").

3. A Framework for Capacity Estimation

To avoid getting lost in the numbers, it's best to follow a structured approach. The following 5-step framework is a reliable way to tackle any estimation problem.

  1. Gather Requirements & State Assumptions: What is the scale of the user base (e.g., Daily Active Users)? What are the usage patterns (read/write ratio)? What kind of data are we storing?
  2. Estimate Traffic (QPS): Calculate the average and peak number of requests your system must handle per second.
  3. Estimate Storage & Bandwidth: Calculate the total amount of data you need to store and the rate at which it moves in and out of your system.
  4. Estimate Memory (Cache): Use the 80/20 rule to estimate how much "hot" data should be kept in a fast in-memory cache.
  5. Estimate Servers: Based on the QPS, calculate the number of servers needed.

Today, we will focus on steps 2 and 3. Let's walk through them with a practical example: designing a simple text-pasting service like Pastebin.

For a guided walkthrough of this process, the following video is an excellent resource. We'll follow its logic closely.

Capacity Planning and Estimation | System Design for Beginners

The presenter, Shiran Afergan, provides a very clear, step-by-step method for these calculations, emphasizing the importance of aggressive rounding to keep things simple.

You can watch the sections on traffic estimation, storage estimation, and bandwidth estimation as you read through the steps below. Pay close attention to how she converts daily users to QPS and daily data generation to total storage needs.

Step 1: Requirements and Assumptions

First, we must state our assumptions. This is the most important step in an interview.

  • Users: 10 million Daily Active Users (DAU).
  • Write Pattern: 10% of users create one paste per day.
  • Read Pattern: The service is read-heavy. Let's assume a 50:1 read-to-write ratio.
  • Data Size: The average paste is 10 KB.
  • Retention: Data is stored for 5 years.

Step 2: Estimate Traffic (QPS)

Now, let's convert those user actions into requests per second.

  • Write Requests per day: 10 million DAU * 10% = 1 million write requests/day.
  • Read Requests per day: 1 million writes/day * 50 (read/write ratio) = 50 million read requests/day.

To get QPS (Queries Per Second), we divide by the number of seconds in a day.
Number of seconds in a day = 24 hours * 60 min/hr * 60 sec/min = 86,400 seconds.
For BOTE, we aggressively round this up to 100,000 () seconds.

  • Average Write QPS: 1,000,000 requests / 100,000 seconds = 10 QPS.
  • Average Read QPS: 50,000,000 requests / 100,000 seconds = 500 QPS.

Traffic is never uniform; it has peaks. A common rule of thumb is that peak traffic is 2x to 3x the average.

  • Peak Write QPS: 10 QPS * 2 = 20 QPS.
  • Peak Read QPS: 500 QPS * 2 = 1000 QPS.

These peak numbers tell us what our system must be able to handle without falling over.

Step 3: Estimate Storage and Bandwidth

Storage Estimation

Here we calculate the total disk space needed.

  1. New data per day: 1 million new pastes/day * 10 KB/paste = 10,000,000 KB/day = 10 GB/day.

    • Calculation check: 1 million () * 10 Kilo () = bytes.
    • Since 1 GB is bytes, this is 10 GB.
  2. Total data over retention period: We need to store data for 5 years.

    • Days in 5 years = 5 years * 365 days/year ≈ 5 * 400 = 2000 days.
    • Total raw data = 10 GB/day * 2000 days = 20,000 GB = 20 TB.
  3. Factor in replication: For durability and availability, data is almost never stored just once. A typical replication factor is 3.

This simple calculation shows how a raw storage estimate is multiplied by a replication factor to get the final required capacity.
*   **Total Storage Needed:** 20 TB * 3 = **60 TB**.

So, we need to provision our database system (like the PostgreSQL or MySQL systems you're familiar with) for at least 60 TB of storage over 5 years.

Bandwidth Estimation

Bandwidth refers to the data transfer rate your network needs to support. We calculate it for incoming (ingress) and outgoing (egress) traffic.

  • Ingress (Write) Bandwidth: This is generated by users uploading new pastes.

    • Write QPS * Average Request Size = 10 QPS * 10 KB = 100 KB/s.
  • Egress (Read) Bandwidth: This is generated by serving pastes to users.

    • Read QPS * Average Response Size = 500 QPS * 10 KB = 5000 KB/s = 5 MB/s.

These numbers are important for network capacity planning and estimating data transfer costs in the cloud.

4. Tips for Success in an Interview

Performing the calculation is only half the battle. Communicating your process is just as important.

Back-of-the-envelope Estimation - System Design

This article from ByteByteGo concludes with some excellent, actionable tips for handling these questions in an interview.

Read the section titled Tips. Pay special attention to the advice on rounding, writing down assumptions, and labeling units. These simple habits prevent costly mistakes and show the interviewer you have a clear thought process.

To summarize the most crucial advice:

  • Communicate your assumptions explicitly. Start by saying, "Let's assume we have X users..." This allows the interviewer to correct you if they have a different scale in mind.
  • Round aggressively. The goal is simplicity and speed. 86,400 becomes 100,000. 365 becomes 400. This is expected.
  • Label everything. Don't just write "10". Write "10 QPS" or "10 GB/day". This avoids confusion for both you and the interviewer.
  • Do a sanity check. After you get a number, ask yourself: "Does this feel right?" If you calculate that a simple app needs more storage than all of Google, you probably made a mistake with your units (e.g., confusing bits and bytes, or MB and GB).

Conclusion

Today, we've transformed the abstract need for "scalability" into concrete numbers. You now have a framework to estimate the load your system will face, which is the first and most critical step in designing an architecture that can handle it.

Key Takeaways:

  • BOTE is about order-of-magnitude estimation, not precision. Its purpose is to quickly assess feasibility and guide high-level design choices.
  • A structured 5-step framework (Requirements, Traffic, Storage/Bandwidth, Cache, Servers) brings order to the estimation process.
  • Key calculations:
    • QPS: Convert Daily Active Users and usage patterns into average and peak requests per second.
    • Storage: Calculate daily data generation, then multiply by the retention period and replication factor.
    • Bandwidth: Calculate ingress (writes) and egress (reads) based on QPS and object size.
  • Communication is critical: Always state your assumptions, use round numbers, and label your units to demonstrate a clear, methodical approach.

In our next lesson, we will explore the CAP theorem. Now that we understand how to quantify the scale of a distributed system, we must grapple with the fundamental trade-offs it imposes on us, specifically regarding data consistency, availability, and partition tolerance.

Can't find a good explanation? Sign up and we'll make it for you

Sign up