Skip to main content
Create your own
Lesson illustration

Scaling Up vs. Scaling Out

Welcome to your next lesson in the "Foundations of Scalable Systems" module. In our previous session, we established the critical distinction between latency ("how fast") and throughput ("how much"). We saw that when a system's throughput can no longer handle the load, we must find ways to increase its capacity.

Today, we'll dive into the two fundamental strategies for doing just that: vertical scaling (scaling up) and horizontal scaling (scaling out). This is one of the most common topics in system design, and having a nuanced understanding of the trade-offs is essential for both building robust systems and succeeding in technical interviews. Your experience architecting backend systems for smaller clients likely involved implicit scaling decisions; this lesson will provide the formal framework to articulate and defend those choices at the scale of large, remote-first companies.

1. Defining the Two Paths to Scale

When your application outgrows its server, you have two choices: get a bigger server or get more servers. This simple idea is the essence of vertical and horizontal scaling.

  • Vertical Scaling (Scaling Up): This involves increasing the resources of a single server. You're making the server more powerful by adding more CPU cores, more RAM, or faster storage (like NVMe SSDs).
  • Horizontal Scaling (Scaling Out): This involves adding more servers to your system and distributing the load among them, usually with a load balancer.

This image provides a clear visual metaphor for the two approaches.

Vertical scaling increases the capacity of a single server, making it "taller." Horizontal scaling adds more servers to the system, making it "wider."

To get a quick, animated overview of these concepts and their initial pros and cons, let's watch a short video.

Vertical Vs Horizontal Scaling: Key Differences You Should Know

The channel ByteByteGo is a highly respected resource for system design concepts. This video offers a concise and clear introduction to our topic.

Watch the entire video. Pay attention to the definitions of vertical scaling and horizontal scaling, and note the key advantages and disadvantages listed for each.

2. A Tale of Two Strategies: Pros, Cons, and Trade-offs

As the video highlighted, neither approach is universally superior. The choice involves a complex set of trade-offs related to cost, complexity, performance, and reliability. The article we'll reference throughout this lesson, from Design Gurus, is specifically framed for system design interviews, which aligns perfectly with your goals.

Let's start by reading the core definitions and the strengths and weaknesses of each approach.

Horizontal vs Vertical Scaling: What System Design Interviews Really Test

This article provides an excellent, pragmatic comparison of the two scaling strategies, focusing on what matters in real-world systems and interviews.

First, read the section on vertical scaling to understand when it works and when it fails. Then, read the section on horizontal scaling for a parallel analysis. As you read, think about how these trade-offs would have applied to the systems you've previously built and managed.

Let's synthesize those points.

Vertical Scaling: The Power of Simplicity

For a team leader like yourself who has managed codebase architecture, the appeal of vertical scaling is its simplicity.

  • Pros:
    • Operational Simplicity: One server means one set of logs, one deployment target, and a single place to debug.
    • No Network Complexity: Communication between processes on the same machine (Inter-Process Communication or IPC) is orders of magnitude faster than network calls between different servers.
    • Strong Consistency: With a single database server, you don't have to worry about the complexities of keeping data synchronized across multiple nodes. This is the default state for the MySQL and PostgreSQL systems you're familiar with.
  • Cons:
    • Hardware Ceiling: You can only buy a machine so big. Eventually, you hit a hard physical limit.
    • Single Point of Failure (SPOF): If your one powerful server goes down, your entire system is offline. High availability is impossible.
    • Cost: High-end enterprise hardware has a non-linear cost curve; doubling the power can triple the price.

Horizontal Scaling: The Power of Distribution

Horizontal scaling is the foundation of modern, cloud-native architecture. It's more complex but offers benefits that are impossible to achieve with a single machine.

  • Pros:
    • High Availability & Resilience: If one of 50 servers fails, the system loses only 2% of its capacity and continues to run.
    • Near-Infinite Scalability: You can theoretically add servers indefinitely to meet demand. This is the primary way to improve a system's throughput.
    • Flexibility: It pairs perfectly with cloud features like auto-scaling, allowing you to match resources to demand precisely.
  • Cons:
    • Architectural Complexity: You now need load balancers, service discovery, distributed configuration, and robust monitoring.
    • Network Latency: Services must communicate over the network, which is inherently slower and less reliable than IPC.
    • Data Consistency Challenges: This is the most significant hurdle. When data is spread across multiple machines, ensuring consistency becomes a major challenge. This directly leads to concepts like eventual consistency, which you wanted to explore, and the trade-offs defined by the CAP and PACELC theorems (which we'll cover soon).

3. Scaling in the Real World: It's a Hybrid Game

The most insightful part of learning system design is seeing how these theoretical concepts are applied—or sometimes ignored—by real companies. The "right" answer isn't a dogmatic choice of one over the other, but a pragmatic application of both.

Let's examine some fascinating case studies.

Horizontal vs Vertical Scaling: What System Design Interviews Really Test

Let's return to the Design Gurus article to see how major tech companies have navigated these scaling decisions.

Read the section on real-world examples. Pay close attention to the stories of Stack Overflow, Twitter, and Instagram. Following that, read the section The Hybrid Approach to see how these strategies are typically combined.

These examples are incredibly instructive:

  • Stack Overflow is a powerful reminder that vertical scaling is not a "beginner" strategy. For their specific workload (read-heavy, well-optimized SQL), a few powerful, vertically-scaled database servers have been the right choice for years.
  • Twitter's "Fail Whale" era is the canonical story of the painful, expensive process of re-architecting a system from a vertically-scaled monolith to a horizontally-scaled distributed system.
  • Instagram's journey is particularly relevant to your background. Instead of abandoning PostgreSQL, they engineered a way to scale it horizontally through sharding. This hybrid approach—vertically scaling each individual database shard while horizontally scaling the number of shards—allowed them to keep the power of a relational database while achieving massive scale.
  • Netflix represents the full commitment to horizontal scaling, where every component is a microservice designed for failure and independent scaling.

The key takeaway is that most large systems are a hybrid. They scale different components differently:

  • Stateless Services (like API servers or web servers using your Go or JS backend) are almost always scaled horizontally. They don't store session data, so adding more servers is easy.
  • Stateful Services (like your MySQL, PostgreSQL, or MongoDB databases) are the challenge. The common pattern is to scale vertically first. When that's not enough, you introduce horizontal scaling for reads via read replicas. Only when write load or total data size becomes unmanageable do you take the much more complex step of sharding.

4. Answering the Interview Question

In a system design interview, your goal is to demonstrate that you understand these trade-offs and can make pragmatic, data-driven decisions. Answering "How would you scale this system?" requires showing a progression of thought.

The Design Gurus article provides excellent rules of thumb and a model answer.

Horizontal vs Vertical Scaling: What System Design Interviews Really Test

Finally, let's focus on how to apply this knowledge in an interview setting.

Read the Decision Thresholds section to get a feel for the orders of magnitude that guide scaling decisions. These are not absolute laws but are invaluable for back-of-the-envelope thinking. Then, carefully study the sample answer in the section How This Shows Up in Interviews.

The sample answer provided is excellent because it:

  1. Starts simple: A single vertically scaled database for the initial launch.
  2. Shows progression: Describes how the architecture evolves with user growth (adding read replicas, then caching).
  3. Identifies a clear trigger for increased complexity: Proposes sharding only when a specific metric (write QPS) exceeds a threshold.
  4. Differentiates components: Scales the stateless application tier horizontally from the start, while treating the stateful database tier differently.

This demonstrates mature engineering judgment—you're not over-engineering a solution for day one but have a clear plan for future growth.

Conclusion

Today we've unpacked one of the most fundamental dichotomies in system design. You now have a robust mental model for comparing vertical and horizontal scaling, not just as technical definitions, but as strategic choices with deep architectural implications.

Key Takeaways:

  • Vertical scaling (up) means a bigger machine. It's simple and fast for single-node operations but has hard limits, high costs, and creates a single point of failure.
  • Horizontal scaling (out) means more machines. It offers resilience and near-limitless scale but introduces significant complexity, especially around data consistency.
  • Statefulness is the deciding factor. Stateless components are easy to scale horizontally. Stateful components (especially databases) are scaled vertically as long as possible before adopting more complex hybrid strategies like read replicas and sharding.
  • Real systems are hybrid. The most effective architectures apply the right scaling strategy to the right component.
  • In interviews, demonstrate pragmatic, progressive thinking. Start simple and evolve the architecture based on specific, articulated capacity needs.

In our next lesson, we will equip ourselves with the tools to justify these scaling decisions with numbers. We'll learn how to perform back-of-the-envelope calculations to estimate the QPS, storage, and bandwidth your system will need to handle, allowing you to predict when you'll hit the limits of one scaling strategy and need to move to the next.

Can't find a good explanation? Sign up and we'll make it for you

Sign up