Hello! Welcome back to our course on designing high-load distributed systems.
In our previous lesson, we focused on making a single Redis instance durable by exploring the RDB and AOF persistence models. We established how to protect your data against server restarts or crashes. However, durability is only one piece of the puzzle. For a system to be truly robust, it must also be highly available and scalable.
Today's lesson directly addresses these needs. Our learning outcome is to configure Redis replication with a master-replica topology. We will move from a single-node setup to a multi-node architecture, which is the foundation for both scaling read traffic and providing fault tolerance. You'll learn how this replication works under the hood, how to configure it, and what architectural trade-offs it entails. This is a fundamental pattern used in virtually all high-load systems that rely on Redis.
1. The Master-Replica Architecture: Scaling Reads and Improving Availability
The most common approach to scaling Redis is to introduce replication. This follows a leader-follower (or master-replica) pattern.
In this architecture:
- One Redis instance is designated as the master. It is the single source of truth and handles all write operations (
SET,INCR,DEL, etc.). - One or more Redis instances are configured as replicas. They maintain a near real-time copy of the master's dataset.
This setup provides two immediate benefits, which are critical in the high-load environments you're familiar with:
- Read Scalability: Client applications can be configured to direct read queries (
GET,SMEMBERS, etc.) to the replicas. This offloads the master, allowing it to dedicate its resources to handling writes and ensuring the system's overall read throughput can be scaled horizontally by simply adding more replicas. - High Availability: A replica serves as a hot standby. If the master node fails, one of the replicas can be promoted to become the new master. While this promotion is a manual process in a basic setup, it drastically reduces downtime compared to restoring a single node from a backup.
The following resource provides a clear, high-level overview of this architecture and its benefits for read-heavy applications.
Replication in Databases with Redis: Architecture for ...
This article, 'Replication in Databases with Redis', introduces the master-replica concept and explains how it facilitates horizontal scaling.
Please read the introductory section 'Horizontal scaling with replicas'. Focus on how write and read operations are distributed and the implications of asynchronous replication on data consistency.
As the article notes, Redis replication is asynchronous by default. The master sends a stream of commands to its replicas but does not wait for their acknowledgment before replying to the client. This prioritizes performance and low latency but introduces a small window where data written to the master might not yet be on the replicas, leading to potential read-after-write inconsistencies.
The official Redis documentation summarizes the key characteristics of this model.
Let's now turn to the official Redis documentation for a more technical summary of replication features.
Please read the section titled 'Important facts about Redis replication'. Pay attention to the non-blocking nature of replication on both the master and replica sides, and the ability to create cascading replicas.
2. How Redis Replication Works
Understanding the synchronization mechanism is key to diagnosing issues and reasoning about the system's behavior during network partitions or node failures. When a replica connects to a master, it uses the PSYNC command to initiate synchronization.
The master maintains a replication ID and an offset. This pair, (replication ID, offset), uniquely identifies a specific version of the dataset's history.
- The replication ID is a pseudo-random string that changes when a replica is promoted to a master or a master restarts without persistence.
- The offset is a counter that increments for every byte of data sent in the replication stream.
Based on the replication ID and offset provided by the connecting replica, the master decides between two synchronization strategies:
-
Partial Resynchronization: If the replica's replication ID matches the master's, and the requested offset is still within the master's in-memory replication backlog buffer, the master sends only the commands the replica missed during its disconnection. This is a highly efficient process for handling brief network interruptions.
-
Full Resynchronization (Full Sync): This occurs if the replica is connecting for the first time, or if its history has diverged too much (i.e., its offset is outside the master's backlog). The process is more intensive:
a. The master performs a background save, creating an RDB snapshot. It buffers all new write commands received during this time.
b. The master transfers the RDB file to the replica.
c. The replica saves the RDB file to disk and then loads it into memory, replacing any old data.
d. The master sends the buffered write commands to the replica to bring it fully up-to-date.
The following reading from the official documentation details this entire process.
This section of the Redis documentation explains the mechanics of synchronization, which is crucial for understanding replication behavior.
Please read the section 'How Redis replication works'. Focus on the roles of the replication ID and offset, and the distinction between a full sync and a partial sync using PSYNC.
3. Configuring a Master-Replica Topology
Now, let's move to the practical implementation. Configuring replication is straightforward. The primary directive is replicaof, which is set in the replica's configuration file or issued as a command.
For a hands-on example, we'll refer to a guide that uses Docker Compose to set up a master and two replicas. This is an excellent way to simulate a multi-node environment on your local machine.
Replication in Databases with Redis: Architecture for ...
This article provides a practical, step-by-step guide to setting up a master-replica environment using Docker Compose.
Please review the 'Implementation' section. You don't need to run the code now, but examine the docker-compose.yml and the redis.conf files. Note the use of the replicaof master 6379 line in the replica's configuration. Also, review the short 'demo' section showing how to verify the setup.
Key Configuration Steps:
-
On the Replica Instance:
- To permanently configure an instance as a replica, add the following line to its
redis.conffile:replicaof <master-ip> <master-port> - To dynamically make a running Redis instance a replica of a master, connect to it with
redis-cliand run:REPLICAOF <master-ip> <master-port>
- To permanently configure an instance as a replica, add the following line to its
-
Authentication (If Master is Password-Protected):
- If the master requires a password (via
requirepass), you must configure the replica to authenticate. Add this to the replica'sredis.conf:masterauth <master-password> - This can also be set dynamically on the replica with
CONFIG SET masterauth <master-password>.
- If the master requires a password (via
-
Verification:
- Once connected, you can verify the replication status on either the master or replica by running
INFO replicationinredis-cli. - On the master, the output will list its connected replicas.
- On the replica, it will show the master's IP and port, and the link status.
- A more structured command is
ROLE, which clearly states if the instance is amasterorslave(the old term for replica) and provides relevant details.
- Once connected, you can verify the replication status on either the master or replica by running
The official documentation provides a concise reference for these configuration directives.
The official documentation provides the definitive reference for the core replication configuration directives.
Quickly scan the 'Configuration' section to see the replicaof directive, and the 'Setting a replica to authenticate to a master' section for masterauth.
4. Operational Considerations
For production systems, simply enabling replication is not enough. You must consider several operational aspects.
-
Read-Only Replicas: By default, replicas are in read-only mode (
replica-read-only yes). This is a critical safety feature that prevents accidental writes to a replica, which would cause its data to diverge from the master. Disabling this is strongly discouraged, as writes to a replica are not propagated to other replicas and will be wiped out upon a full resynchronization. -
Data Safety Guarantees: Because replication is asynchronous, a write acknowledged by the master could be lost if the master crashes before the write propagates to any replica. To mitigate this, Redis offers a "best-effort" data safety mechanism. You can configure the master to accept writes only if it has a minimum number of replicas connected with a lag of no more than a specified number of seconds.
min-replicas-to-write <number>min-replicas-max-lag <seconds>
This configuration forces a trade-off: it increases durability guarantees at the cost of availability (the master will reject writes if the conditions aren't met).
-
Replication in Containerized Environments (Docker/NAT): When using port mapping, a master might see a replica connecting from an internal Docker IP address. To ensure the master reports the correct, publicly accessible address for its replicas, you can use these directives on the replica:
replica-announce-ip <ip-address>replica-announce-port <port>
These advanced settings are crucial for running a resilient Redis setup in a real-world production environment.
Conclusion
In this lesson, we've established the foundation for building scalable and highly available Redis deployments.
Key Takeaways:
- Master-Replica Architecture: This model is the standard for scaling Redis reads and providing a basis for high availability. Writes go to the master, and reads are distributed across replicas.
- Asynchronous Replication: Prioritizes performance but introduces a small window of potential data loss and inconsistency.
- Synchronization Mechanisms: Redis uses an efficient partial resynchronization for minor disruptions and a more intensive full resynchronization for initial setup or major disruptions.
- Core Configuration: The
replicaofdirective is the key to establishing the relationship, withmasterauthbeing essential for secured masters. - Operational Trade-offs: Features like
min-replicas-to-writeallow you to tune the system's behavior, balancing availability and data safety to meet your application's specific requirements.
Preview of the Next Lesson:
Our current master-replica setup still has a major weakness: if the master fails, a manual intervention is required to promote a replica. This is not ideal for high-load systems that demand automatic recovery. In our next lesson, we will solve this by learning how to set up Redis Sentinel for automatic failover and high availability, turning our replicated setup into a self-healing system.