Skip to main content
Create your own
Lesson illustration

Async vs. Sync Replication: Trade-offs

Hello! In our last lesson, we explored the PACELC theorem, which highlights the fundamental trade-off between latency and consistency that distributed systems face during normal operation. We established that this choice exists. Now, we'll dive into how this choice is implemented at the database level.

This lesson focuses on database replication, the core mechanism for maintaining multiple copies of your data. Our goal is to dissect the trade-offs between the two primary modes: synchronous and asynchronous replication. Understanding these strategies is not just theoretical; it's a practical necessity for designing resilient and performant systems, and a common point of discussion in senior software engineering interviews.

1. Why Replicate Data?

Before comparing replication modes, let's briefly establish why we need it in the first place. A single database instance is a single point of failure. If it goes down, your entire application can become unavailable. It can also become a performance bottleneck as user traffic grows. Replication addresses these issues by creating copies (replicas) of your database.

The following video provides a great overview of the fundamental motivations for replication.

Database Replication | Synchronous vs Asynchronous | System Design Interview Question | Code Decode

This video from Code Decode, "Database Replication," clearly explains the business and technical drivers behind replication.

Watch from the beginning to the introduction. Focus on the three key benefits the presenter outlines: High Availability: Protecting against server crashes. Increased Performance: Distributing the load, especially for read queries. Reduced Latency: Placing data geographically closer to users (though we'll focus more on the first two).

With the "why" established, let's look at the "how." The way data is copied from the primary (or leader) database to its replicas dictates the guarantees your system can provide. This leads us to the core choice between synchronous and asynchronous replication.

The diagram below illustrates the fundamental difference in the data flow for these two modes.

This diagram shows the sequence of operations for synchronous and asynchronous replication. For synchronous, the leader waits for success confirmation from the follower (step 4) before confirming to the client (step 5). For asynchronous, the leader confirms success to the client immediately (step 6) after sending the data change to the follower (step 3).

Let's break down each of these approaches.

2. Synchronous Replication: The Path of Consistency

In synchronous replication, when a client sends a write request (e.g., INSERT, UPDATE), the primary database performs the write and then forwards the change to its synchronous replicas. The crucial step is that the primary waits for confirmation from at least one (and sometimes all) of those replicas before it sends a "success" acknowledgement back to the client.

This directly maps to the PACELC framework we discussed. By waiting for replicas, you are explicitly choosing Consistency over Latency (EC). You accept higher write latency to guarantee that the data is safely stored on multiple nodes before you confirm the operation.

Database Replication Modes (Async vs Sync) - DEV Community

The "The Synchronous Tango" section of this DEV Community article provides a clear, detailed breakdown of synchronous replication.

Read the entire section, from The Synchronous Tango. Pay close attention to the list of advantages, disadvantages, and the specific use cases provided. Note the conceptual PostgreSQL example, which should feel familiar given your experience.

As the article highlights, the trade-offs are clear:

  • Advantage: Strong Consistency & Durability. If the primary server crashes immediately after confirming a write, the data is not lost because it's guaranteed to exist on at least one replica. This is critical for systems where data loss is unacceptable.
  • Disadvantage: Higher Latency & Reduced Availability. The write operation now includes the network round-trip time to the slowest replica. This increases the overall response time for the client. Furthermore, if a synchronous replica fails, the primary might be unable to commit any new writes, effectively making the write path of your system unavailable.

A classic interview question would be: "When would you insist on synchronous replication?" Your answer should revolve around use cases where data integrity is paramount:

  • Financial transactions: A confirmed payment must never disappear.
  • Inventory management: An e-commerce site cannot afford to sell the same last item to two different customers.
  • Authentication systems: A password change must be reflected everywhere immediately.

3. Asynchronous Replication: The Path of Speed

In asynchronous replication, the primary database writes the data locally, sends a "success" acknowledgement back to the client, and then sends the data changes to its replicas in the background. The primary does not wait for confirmation from the replicas.

This is the other side of the PACELC coin: choosing Latency over Consistency (EL). The system prioritizes responding to the client as fast as possible, accepting that there will be a small window of time where the replicas are not up-to-date.

Database Replication Modes (Async vs Sync) - DEV Community

Now, let's explore the alternative in the same article.

Read the section titled The Asynchronous Waltz. Again, focus on the pros, cons, and use cases. The conceptual MySQL example illustrates the standard configuration for this mode.

The trade-offs are the inverse of the synchronous approach:

  • Advantage: Lower Latency & High Throughput. Writes are fast because the client only waits for the primary to complete its local write. The primary is also decoupled from the health of the replicas, so a slow or failed replica won't block new writes.
  • Disadvantage: Potential for Data Loss. This is the critical risk. If the primary crashes after confirming a write to the client but before the data has been sent to any replica, that write is permanently lost. The client thinks the data was saved, but upon failover to a replica, the data is gone. This is sometimes called a "ghost write."

Asynchronous replication is the right choice when performance is key and brief staleness or a small amount of data loss is tolerable:

  • Analytics and logging: Losing a few tracking events is often acceptable.
  • Social media feeds: Seeing a new post a few seconds late is not a critical failure.
  • Content Management Systems (CMS): An updated article appearing on replicas with a slight delay is fine.

4. Comparing the Modes and Real-World Consequences

To solidify your understanding, let's look at a direct comparison and a cautionary tale.

Database Replication Modes (Async vs Sync) - DEV Community

The article provides a handy summary table that's perfect for quick review.

Focus on the comparison table in the "The Balancing Act" section. This is an excellent cheat sheet for interviews.

This table is a great summary, but stories of failure often drive the point home more effectively. The choice of replication mode isn't just a configuration detail; it has real, tangible consequences for the business.

Distributed Systems Trade-offs: CAP, PACELC & Replication - Tarento

This article from Tarento connects replication directly to system design decisions and includes a powerful example.

Read the section Replication: Synchronous vs. Asynchronous. Pay special attention to the "Failure story".

The story of the payments team that lost "confirmed" orders because they used asynchronous replication is a classic example of misaligning technical choices with business requirements. This is exactly the kind of architectural debt you want to avoid as a system designer.

5. Hybrid Approaches: Getting the Best of Both Worlds

Fortunately, the choice isn't always a strict binary. Many modern systems, including PostgreSQL, allow for more nuanced, hybrid configurations. You can have a mix of synchronous and asynchronous replicas for the same primary database.

A PostgreSQL primary server replicating to four standby servers. Replicas S1, S2, and S3 are synchronous, meaning the primary waits for their confirmation. Replica S4 is asynchronous. In this scenario, replica S1 has failed, but the primary can still commit writes as long as S2 and S3 are healthy (depending on the exact configuration).

This hybrid model is incredibly powerful. You might configure it this way:

  • Synchronous Replicas (S1, S2, S3): These are located in the same or adjacent data centers. They are used for high-availability failover. If the primary fails, one of these can be promoted with zero data loss.
  • Asynchronous Replica (S4): This replica might be in a different geographical region for disaster recovery, or it could be used to serve traffic for an analytics dashboard. The slight data lag is acceptable for these use cases.

PostgreSQL also allows for "quorum-based" synchronous replication, where you can specify that the primary must wait for confirmation from any N of a list of synchronous standbys, adding resilience if one of them fails.

The following video discusses these replication modes specifically in the context of PostgreSQL.

PostgreSQL Streaming Replication Tutorial

This tutorial on PostgreSQL Streaming Replication directly addresses how these modes are implemented.

Watch the segments on asynchronous replication, then synchronous replication. Finally, watch the short concluding piece on mixing the two modes, which reinforces the concept of the hybrid approach.

Conclusion

In this lesson, we connected the theoretical trade-off of Consistency vs. Latency from PACELC to the concrete implementation of database replication. You can now confidently describe the mechanisms and consequences of choosing one mode over the other.

Key Takeaways:

  • Synchronous Replication: Prioritizes consistency and durability at the cost of higher latency and potentially lower availability. It's the choice for critical data where loss is unacceptable.
  • Asynchronous Replication: Prioritizes low latency and high throughput at the risk of data loss in a failure scenario. It's suitable for less critical data or workloads that can tolerate eventual consistency.
  • The Choice is a Trade-off: The "right" mode depends entirely on the business requirements for a specific piece of data. As the failure story showed, a mismatch can have severe consequences.
  • Hybrid Models are Common: Real-world systems often use a mix of synchronous replicas for high-availability failover and asynchronous replicas for disaster recovery or read scaling.

You now have the foundational knowledge of why and how data is replicated. In our next lesson, we will put this into practice by focusing on a very common use case: implementing read replicas with PostgreSQL to scale out a system handling read-heavy traffic.

Can't find a good explanation? Sign up and we'll make it for you

Sign up