Skip to main content
Create your own

MongoDB Consistency: Read/Write Concerns

Hello! Welcome to the first lesson in our module on MongoDB.

In this lesson, we will explore MongoDB's consistency model. Your goal is to move beyond theoretical knowledge and understand the practical levers available in real-world systems. MongoDB's approach is a prime example of this, offering tunable consistency that allows architects to make deliberate trade-offs between performance, durability, and data recency.

This topic directly addresses your interest in the consistency models of different databases. Unlike the rigid ACID guarantees of traditional relational databases (which we'll cover in the PostgreSQL module), MongoDB's flexibility is a key reason for its adoption in high-load, distributed environments.

By the end of this lesson, you will be able to explain MongoDB's consistency model and configure read and write concerns for different consistency requirements.

1. Replication and the Consistency Spectrum

MongoDB achieves high availability and data redundancy through replica sets. A replica set is a group of mongod instances that maintain the same data set. One node is the primary, which receives all write operations. The other nodes are secondaries, which replicate the primary's operation log (oplog) and apply the operations to their own data sets.

This replication model is the foundation of MongoDB's consistency controls. When a client writes data, when does the database acknowledge it? When a client reads data, which version of the data does it see? The answers are not fixed; you control them using Write Concerns and Read Concerns.

Let's look at a visual representation of how a write concern can change the behavior of a write operation.

This diagram illustrates the difference between a default write concern and a more durable one. On the left, the application receives a response as soon as the primary has processed the write. On the right, with `writeConcern: 2`, the application must wait for the write to be replicated to at least one secondary, increasing durability at the cost of higher latency.

This diagram highlights the fundamental trade-off you'll be managing. Let's break down the tools you have to control it.

2. Ensuring Write Durability: Write Concern

A write concern defines the level of acknowledgment requested from MongoDB for write operations. It's how you tell the database how durable a write needs to be before it's considered successful.

To understand the components, let's turn to the official documentation.

Write Concern - Database Manual - MongoDB Docs

To start, let's dive into the specifics of controlling write durability. The official MongoDB documentation provides a comprehensive guide. Please read this short section to learn about the three key components of a write concern specification.

Please read the 'Write Concern Specification' section. Focus on understanding the purpose of the w, j, and wtimeout options.

As you just read, a write concern is a combination of three parameters:

  • w: The core of the write concern. It specifies the number of mongod instances that must acknowledge the write.

    • w: 0: "Fire-and-forget". The driver doesn't wait for any acknowledgment. This offers the lowest latency but provides no guarantee the write was received. Useful for non-critical data like logging at extreme volumes.
    • w: 1: Acknowledgment from the primary node only. This was the default before MongoDB 5.0. It's fast but leaves the write vulnerable to being rolled back if the primary fails before replicating it.
    • w: "majority": Acknowledgment from a calculated majority of voting members in the replica set. This is the current default and the key to durability. A write acknowledged by a majority is guaranteed not to be rolled back, as any new primary elected must have seen this write.
    • w: <number>: Acknowledgment from a specific number of nodes (primary + number-1 secondaries).
  • j: The journaling option. If j: true, the acknowledgment is sent only after the write has been committed to the on-disk journal on the specified number of nodes. This ensures durability even if the mongod processes crash. For w: "majority", journaling is typically enabled by default on the voting members, providing a very strong durability guarantee.

  • wtimeout: A crucial safety net. This is a time limit (in milliseconds) for the write concern to be satisfied. If the required number of nodes don't acknowledge the write within this period (e.g., due to network issues or unavailable secondaries), the operation fails with a write concern error. This prevents your application from blocking indefinitely and is a key component of building resilient systems.

The Nuance of "Majority"

The concept of "majority" is critical, but its calculation isn't always straightforward, especially in topologies with arbiters (non-data-bearing voting nodes).

Write Concern - Database Manual - MongoDB Docs

The 'majority' calculation has some important subtleties. Please read the section that explains this, as it has direct implications for cluster architecture.

Read the section titled 'Calculating Majority for Write Concern'. Pay close attention to the examples, particularly the Primary-Secondary-Arbiter (P-S-A) case and the warning associated with it.

The key takeaway is that in a Primary-Secondary-Arbiter (P-S-A) architecture, a w: "majority" write requires acknowledgment from all data-bearing nodes (the primary and the single secondary). If the secondary is down, no majority write can ever succeed. This makes the architecture brittle. The recommendation is to use an odd number of data-bearing nodes (e.g., P-S-S) whenever possible to avoid this availability pitfall.

3. Controlling Data Views: Read Concern

Read concern is the other side of the coin. It allows you to control the isolation and consistency of the data you read. You can specify whether a read should return the most recent data on a single node (which might be rolled back) or a version of the data that is durable across the replica set.

Let's review the different levels available.

Read Concern - Database Manual - MongoDB Docs

Now let's look at how we control the consistency of data we read. The 'Read Concern' documentation details the different levels of isolation you can request.

Please read the 'Read Concern Levels' section. Focus on the guarantees provided by 'local', 'majority', and 'linearizable'. Note the differences in performance, availability, and typical use cases for each.

Here is a summary of the most important read concern levels:

  • local: This is the default level. It returns data from the perspective of the mongod instance receiving the query. There's no guarantee that this data has been written to a majority of replicas, so it could be rolled back. This offers high performance but weak consistency.
  • available: The most lenient level. For sharded clusters, it can return data that is "orphaned" during chunk migrations. It offers the lowest latency but should be used with caution.
  • majority: This is the counterpart to w: "majority". It returns data that has been acknowledged by a majority of the replica set members. This data is durable and guaranteed not to be rolled back. This is the most commonly used level for reads that require consistency.
  • linearizable: This provides the strongest consistency guarantee. When you issue a read with this concern against the primary, it ensures that the data returned reflects all successful majority-acknowledged writes that completed before the read operation started. It effectively makes the system behave as if there is only a single copy of the data. This guarantee comes at a significant performance cost, as the primary must confirm its status with a majority of secondaries for each such read. It is only available for reads on a primary that uniquely identify a single document.
  • snapshot: Available only within multi-document transactions, this concern provides a consistent point-in-time snapshot of the data across shards.

4. Practical Scenarios: Combining Read and Write Concerns

The true power of MongoDB's model comes from combining read and write concerns to meet specific application requirements.

Scenario 1: High-Throughput, Non-Critical Data

  • Use Case: Storing application logs, tracking user clicks for analytics.
  • Requirement: High write throughput, low latency. Occasional data loss is acceptable.
  • Configuration:
    • Write Concern: {w: 1, j: false} or even {w: 0}. This minimizes write latency by acknowledging writes quickly or not at all.
    • Read Concern: local. Reads for analysis can be performed with the default concern, as up-to-the-millisecond accuracy is not required.

Scenario 2: Critical Financial or E-commerce Data

  • Use Case: Processing a payment, updating user order status.
  • Requirement: High durability and strong consistency. Data loss is unacceptable.
  • Configuration:
    • Write Concern: {w: "majority", j: true}. This ensures the write is durable on a majority of nodes and written to their journals, making it safe from both rollbacks and crashes.
    • Read Concern: majority. When reading the status of the payment or order, this ensures you are seeing a committed version of the data that will not be rolled back.

Scenario 3: "Read Your Own Writes"

  • Use Case: A user updates their profile and immediately reloads the page to see the change.
  • Requirement: The read following a write must see the result of that write.
  • Configuration:
    • Write Concern: {w: "majority"}.
    • Read Concern: majority.
    • Explanation: Using w: "majority" ensures the write is durable and will not be rolled back. The subsequent read with readConcern: "majority" guarantees that the query engine will only consider data that has been committed by a majority, which by definition includes your preceding write. Using a weaker write concern (w:1) or read concern (local) could result in the read being directed to a secondary that has not yet replicated the write, leading to a stale read.

To help you conceptualize how these settings map to broader consistency models, the following table from Microsoft Azure documentation is quite useful. It compares native MongoDB to the Cosmos DB API, but the mapping on the left side is a great summary.

This table maps MongoDB's native read/write concern combinations to formal consistency models. For example, combining a `MAJORITY` write concern with a `LINEARIZABLE` read concern on the primary achieves "Strong" (Linearizable) consistency. A `MAJORITY` write with a `MAJORITY` read achieves "Consistent Prefix" and prevents reading rolled-back data.

Conclusion

In this lesson, we've dissected MongoDB's tunable consistency model. You now have a framework for making informed decisions about data durability and isolation in your applications.

Key Takeaways:

  • MongoDB's consistency is not one-size-fits-all; it's configured on a per-operation basis using Read and Write Concerns.
  • Write Concern (w, j, wtimeout) controls the durability of writes. w: "majority" is the modern default and the foundation for building consistent systems, as it prevents data rollbacks.
  • Read Concern (local, majority, linearizable) controls the isolation of reads, allowing you to choose between reading the latest (but possibly transient) data and reading majority-committed (durable) data.
  • The combination of w: "majority" and readConcern: "majority" provides strong consistency guarantees suitable for most critical applications, without the significant performance overhead of linearizable.

In our next lesson, "Set up MongoDB replica sets with primary-secondary replication and automatic failover," we will put this theory into practice. We'll explore the operational side of the replica set, the mechanism that underpins these consistency guarantees and provides the high availability MongoDB is known for.

Can't find a good explanation? Sign up and we'll make it for you

Sign up