Skip to main content
Create your own
Lesson illustration

Understanding the CAP Theorem

In our previous lesson, we learned how to perform back-of-the-envelope calculations to estimate the scale a system needs to handle. Those estimations often reveal the necessity of distributing our services across multiple machines. Now that you can quantify why you need a distributed system, we must confront the fundamental trade-offs inherent in building one.

This lesson introduces the CAP theorem, a cornerstone of distributed systems theory. Understanding it is non-negotiable for designing scalable applications and a frequent topic in system design interviews. Our goal is to analyze the theorem's three components—Consistency, Availability, and Partition tolerance—and explore the critical trade-offs they force upon us, particularly when choosing a distributed data store.

1. Defining the CAP Theorem

The CAP theorem states that in a distributed system, it is impossible to simultaneously guarantee all three of the following properties:

  • Consistency (C): All nodes in the cluster see the same data at the same time. Every read operation returns the result of the most recent successful write.
  • Availability (A): Every request receives a (non-error) response, without the guarantee that it contains the most recent write. The system is always up for reads and writes.
  • Partition Tolerance (P): The system continues to operate despite an arbitrary number of messages being dropped (or delayed) by the network between nodes. In essence, the system can sustain a network "partition" that splits the cluster into two or more groups of nodes that cannot communicate with each other.

To get a quick and clear overview, let's start with a short video from ByteByteGo.

CAP Theorem Simplified

This video provides a concise definition of the CAP theorem and its three components.

Watch the first part of the video, from the beginning to the explanation of the core trade-off. This will give you the foundational vocabulary we'll build upon.

The most crucial takeaway is that modern, large-scale systems are built on unreliable networks and span multiple data centers. Network partitions are not a hypothetical risk; they are an operational certainty. Because of this, Partition Tolerance (P) is not optional. A system that isn't partition-tolerant is not a truly distributed system.

This reality simplifies the theorem for practical purposes: during a network partition, you must choose between Consistency and Availability.

This table summarizes the primary use cases and trade-offs for Consistency (C), Availability (A), and the resulting system types (CP and AP). Notice that 'CA' is listed as rare, as it presumes no network partitions, an unsafe assumption in distributed systems.

2. The Core Trade-Off: CP vs. AP

When a network partition happens, a distributed system must make a difficult choice. Imagine your database is replicated across two data centers, one in the US and one in Europe, and the network link between them fails.

  1. Choose Consistency (a CP system): To ensure all clients see the same data, you might have to stop accepting writes in one data center (or even both) to prevent them from diverging. If a client tries to read data, you might have to return an error if you can't be sure you are providing the absolute latest version. In this scenario, you sacrifice availability to maintain consistency.

  2. Choose Availability (an AP system): You could allow both data centers to continue accepting reads and writes. The system remains available to all users. However, their data will now diverge. A user in Europe might not see a write made by a user in the US until the partition is resolved. Here, you sacrifice consistency to maintain availability.

The following video offers excellent, practical examples that make this trade-off concrete.

CAP Theorem in System Design Interviews

This video from Hello Interview is specifically tailored for system design interviews. It clearly frames the C vs. A choice and connects it to real-world applications and technology choices.

First, watch the section from the introduction to understand why Partition Tolerance is considered a given. Then, watch the detailed example from the server replication scenario, which categorizes systems like ticket booking, inventory management, and social media feeds as either CP or AP. Finally, watch the segment on design implications, which discusses how this choice influences your selection of databases and architecture.

As the video explains, the right choice depends entirely on the business requirements:

  • CP (Consistency/Partition Tolerance): Chosen when data correctness is paramount. Incorrect data is worse than temporary downtime.
    • Examples: Financial systems (avoiding double spending), e-commerce inventory management, airline booking systems.
    • Database Leaning: Your experience with MySQL and PostgreSQL in single-master setups aligns with this model. They are often used as the source of truth, prioritizing consistent data.
  • AP (Availability/Partition Tolerance): Chosen when uninterrupted service is more critical than having perfectly up-to-date data.
    • Examples: Social media feeds (seeing a slightly old post is acceptable), content delivery networks, user comment sections.
    • Database Leaning: Some NoSQL databases, like Cassandra and DynamoDB, are famously designed for high availability. Your experience with MongoDB can fit here, as it can be configured for higher availability at the cost of immediate consistency across all replicas.

3. A More Nuanced View: Beyond the CP/AP Labels

The "CP vs. AP" framework is an excellent starting point, but for a senior engineering role, you need a more nuanced understanding. The formal definitions in the CAP theorem are stricter than many people realize, and most real-world systems don't fit neatly into these buckets.

This is a topic brilliantly explored by Martin Kleppmann, a leading researcher in distributed systems.

Please stop calling databases CP or AP

This blog post is a classic critique of the oversimplification of the CAP theorem. It argues that the formal definitions of Consistency and Availability are so precise that many real-world databases don't strictly satisfy them. This perspective is what will set you apart in a senior interview.

First, read the sections What CAP Means. Focus on the strict definitions of C (which is linearizability), A, and P. Next, read the section on Linearizability. The football world cup example is a fantastic, intuitive way to understand this strong consistency model. Then, read the proof walkthrough in The Trade-off. This solidifies the core logic. Finally, and most importantly, read The Reality of Databases. This section connects directly to your experience, explaining why a standard replicated PostgreSQL or MongoDB setup is often neither strictly CAP-Consistent nor CAP-Available.

Let's unpack the key insights from Kleppmann's article:

  • Consistency in CAP means Linearizability: This is a very strong guarantee. It means that if operation B starts after operation A completes, operation B must see the state of the system as it was after A finished, or newer. A typical asynchronous read replica in a PostgreSQL setup violates this, because it might be lagging behind the primary.
  • Availability in CAP is also strict: It means every non-failing node can process a request. In a single-leader database system, if a client is partitioned from the leader, it cannot perform a write. Even if other follower nodes are up, they can't accept the write, so the system is not CAP-Available.
  • Most Systems are "P" with Nuances: A standard PostgreSQL or MySQL database with asynchronous replication is technically neither CP nor AP. It's not CAP-Consistent due to replication lag, and it's not CAP-Available because only the master accepts writes. Similarly, MongoDB's behavior depends heavily on its configuration ("read concern" and "write concern").

The point is not to be pedantic. The point is to recognize that CAP is a powerful model for framing the conversation, but real systems involve a spectrum of trade-offs. In an interview, showing you understand this complexity demonstrates deep expertise.

4. How to Discuss CAP in a System Design Interview

Armed with this nuanced understanding, you can structure a compelling discussion. Use the CAP theorem as a framework to justify your design choices based on user and business needs.

The visual reference below is a great summary of how these concepts fit together.

This infographic brings together the CAP triangle, demonstrates the choice during a network partition, introduces the PACELC theorem (which we'll cover next), and classifies common databases according to their typical behavior.

Here’s a practical guide to discussing CAP in an interview, drawing from the System Design Handbook article.

CAP Theorem in System Design: A Complete Interview ...

This guide provides excellent, actionable advice on how to communicate your understanding of CAP in a high-pressure interview setting.

Focus on these three sections: Common Misconceptions: Internalize these points to avoid common traps. The key is to emphasize that the trade-off is forced during a partition. Applying CAP in Interviews: This shows you how to use CAP as a tool to guide your design narrative. Structuring Your Explanation: Pay close attention to the "Weak Explanation" vs. "Strong Explanation" table.

To synthesize for an interview:

  1. Acknowledge P is a given: Start by stating that for any distributed system, you must design for network partitions.
  2. Frame the choice as C vs. A: Explain that the core decision is whether to prioritize consistency or availability when a partition occurs.
  3. Justify your choice with business needs: Connect your choice to the specific feature you are designing. "For the payment processing part of this system, we must prioritize consistency (CP) to avoid financial errors. However, for the user profile service, we can prioritize availability (AP) and tolerate eventual consistency, as showing a slightly outdated profile picture is not a critical failure."
  4. Discuss the implementation: Your choice of CP or AP will inform your choice of database, replication strategy, and communication patterns. For example, "Since we're choosing CP for our inventory service, we might use a relational database like PostgreSQL with synchronous replication, or a system like Zookeeper for distributed locks. For the AP-focused recommendation engine, we could use a database like Cassandra and design for eventual consistency."

Conclusion

In this lesson, we dissected the CAP theorem, moving from its basic definition to the practical trade-offs it imposes on distributed system design. You are now equipped not just to recite the theorem, but to analyze its implications with the nuance expected of a senior engineer.

Key Takeaways:

  • The CAP theorem states that a distributed system can only provide two of three guarantees: Consistency, Availability, and Partition Tolerance.
  • In modern systems, Partition Tolerance (P) is a necessity, forcing a choice between Consistency (CP) and Availability (AP) during a network failure.
  • The choice between CP and AP is a business decision. CP is for systems where correctness is critical (e.g., finance). AP is for systems where being responsive is critical (e.g., social media).
  • The formal definitions of "Consistency" (linearizability) and "Availability" are very strict. Many real-world systems, including those you've used like PostgreSQL and MongoDB, have behaviors that don't fit neatly into the CP/AP buckets. Understanding this nuance is a sign of seniority.
  • In an interview, use CAP as a framework to explain why you are making certain architectural decisions based on user and business impact.

We've focused on the trade-offs that arise during a network partition. But systems operate normally most of the time. Even then, there are trade-offs to be made, primarily between latency and consistency. This brings us to the PACELC theorem, an extension of CAP that we will explore in our next lesson.

Can't find a good explanation? Sign up and we'll make it for you

Sign up