Hello! Welcome back to our module on Cassandra.
In our previous lesson, we explored how Cassandra uses consistent hashing and virtual nodes to partition data across the cluster. This process determines the primary node responsible for any given piece of data. Now, we'll build directly on that foundation to address the critical needs of fault tolerance and data locality.
Today's lesson addresses the learning outcome: Configure Cassandra replication strategy and replication factor across datacenters.
We will focus on the practical steps and strategic decisions involved in making your data resilient and globally available. We'll cover:
- The roles of Replication Factor and Replication Strategy.
- How to configure a Cassandra cluster to be "topology-aware" using snitches and property files, a prerequisite for multi-datacenter replication.
- Using the
NetworkTopologyStrategyto precisely control how many replicas of your data are stored in each datacenter.
This is a cornerstone of designing the kind of robust, high-availability distributed systems you've worked on, as it directly governs disaster recovery capabilities and read/write latency for a global user base.
1. Replication Fundamentals: Strategy and Factor
While partitioning tells Cassandra which node is the primary owner of a piece of data, replication determines which other nodes will store copies, or replicas, of that data. This is configured on a per-keyspace basis.
Two key settings govern this process:
-
Replication Factor (RF): This is an integer that defines the total number of copies of each row that the cluster should maintain. A replication factor of
1means no redundancy (a single point of failure), while a common production value is3. -
Replication Strategy: This is the algorithm Cassandra uses to place the replicas on the nodes in the cluster. There are two main strategies:
SimpleStrategy: This is a basic strategy intended for single-datacenter deployments and development environments. It places replicas on consecutive nodes along the token ring, without regard for their physical location (e.g., which rack or datacenter they are in). It is not suitable for production or multi-datacenter setups.NetworkTopologyStrategy: This is the recommended strategy for all production deployments. It is "topology-aware," meaning it understands how nodes are grouped into datacenters and racks. It allows you to specify the replication factor independently for each datacenter, which is the primary focus of this lesson.
To use NetworkTopologyStrategy, we must first tell Cassandra about our cluster's physical layout.
2. Making the Cluster Topology-Aware
Before you can intelligently place replicas across datacenters, every node in the cluster needs to know its own location and the location of all other nodes. This is the job of the snitch.
The most common and recommended snitch for production is the GossipingPropertyFileSnitch. It works in two parts:
- It reads a local configuration file on each node to determine that node's datacenter and rack.
- It uses the gossip protocol (which Cassandra already uses for failure detection) to propagate this topology information to all other nodes in the cluster.
Configuring this involves editing two key files on each node.
How to Install Cassandra Across Multiple Data Centers
The following guide from Linode provides a detailed, practical walkthrough of configuring a Cassandra cluster across multiple datacenters. We will focus on the specific file changes required to define the cluster's topology.
Please read the section titled 'How to Configure Cassandra to Run in Multiple Data Centers'. Focus on steps 7 through 14. Pay close attention to the changes made in cassandra.yaml (specifically endpoint_snitch) and the purpose of the cassandra-rackdc.properties file.
Let's summarize the critical configuration steps from the reading:
-
In
cassandra.yaml:cluster_name: Must be identical on every node in the entire cluster.seed_provider: The list of seed nodes. For a multi-DC cluster, the best practice is to include a few nodes from each datacenter. The article also notes an important convention: list the local DC's seeds first.endpoint_snitch: This must be set toGossipingPropertyFileSnitchto enable topology awareness.
-
In
cassandra-rackdc.properties:- This file is very simple and contains just two properties:
dcandrack. dc: You assign a name for the datacenter the node belongs to (e.g.,dc=londonordc=singaporein the example).rack: You assign a name for the failure domain within the datacenter (e.g.,rack=rack1). Racks are important becauseNetworkTopologyStrategywill attempt to place replicas in different racks within the same datacenter to protect against rack-level failures (like a power outage to a single server rack or a failed top-of-rack switch).
- This file is very simple and contains just two properties:
With these settings applied and Cassandra restarted, each node will gossip its location, and the entire cluster will build a map of the network topology.
3. Defining a Multi-DC Replication Strategy
Now that the cluster is topology-aware, we can create a keyspace that leverages this information. This is done using a CREATE KEYSPACE command in CQL.
The NetworkTopologyStrategy allows you to specify exactly how many replicas you want in each datacenter.
How to Install Cassandra Across Multiple Data Centers
This next section of the Linode guide demonstrates how to create a keyspace using NetworkTopologyStrategy.
Please read the section 'How to Add Tables and Data to Cassandra', focusing on step 2, which shows the CREATE KEYSPACE command. Note how the replication factor is defined per-datacenter.
As the article shows, the syntax is clear and powerful:
CREATE KEYSPACE store WITH REPLICATION =
{
'class' : 'NetworkTopologyStrategy',
'london' : 2,
'singapore' : 2
};
In this example:
- The keyspace
storewill useNetworkTopologyStrategy. - For any data written to this keyspace, Cassandra will ensure 2 replicas are stored on nodes in the
londondatacenter. - It will also ensure 2 replicas are stored on nodes in the
singaporedatacenter. - The total replication factor for any piece of data is
2 + 2 = 4.
Cassandra uses this information along with the token ring to place replicas. It first identifies the primary node for a given partition key (as we saw in the last lesson). Then, it walks the ring clockwise to find additional nodes, skipping nodes that are not in the desired datacenter or are in a rack that already has a replica, until its placement obligations are met for all datacenters.
The image below visualizes this concept.

4. Practical Application and Trade-offs
Your experience in fintech and payment systems highlights the importance of this configuration. Let's consider a practical scenario.
Imagine you are designing a global payment processing system with datacenters in us-east-1 and eu-central-1.
-
User Data (
profileskeyspace): User data must be highly available in both regions for low-latency authentication. You might choose a symmetric replication:CREATE KEYSPACE profiles WITH REPLICATION = { 'class' : 'NetworkTopologyStrategy', 'us-east-1' : 3, 'eu-central-1' : 3 };This ensures that a full datacenter outage in either region will not result in data loss, and local reads can be served quickly to users in both North America and Europe.
-
Regional Logs (
analytics_eukeyspace): You might have analytics data that is only relevant to the European business unit and subject to data residency regulations like GDPR. In this case, you can configure replication to a single DC:CREATE KEYSPACE analytics_eu WITH REPLICATION = { 'class' : 'NetworkTopologyStrategy', 'eu-central-1' : 3 };Data in this keyspace will only ever be stored on nodes physically located in the
eu-central-1datacenter.
This per-keyspace control over data placement is a powerful tool for balancing availability, performance, cost, and compliance requirements in a distributed environment.
Conclusion
In this lesson, we've moved from partitioning to replication, focusing on the practical steps to configure a multi-datacenter Cassandra cluster.
Key Takeaways:
- Replication is configured per-keyspace using a Replication Strategy and Replication Factor.
NetworkTopologyStrategyis the required choice for production and multi-DC clusters.- To use it, the cluster must be made topology-aware by setting the
endpoint_snitchtoGossipingPropertyFileSnitchincassandra.yamland defining thedcandrackfor each node incassandra-rackdc.properties. - With
NetworkTopologyStrategy, you define the replication factor for each datacenter individually, giving you precise control over data distribution for availability, performance, and compliance.
Preview of the Next Lesson:
We have now defined how data is partitioned (Partitioner), where replicas are placed (ReplicationStrategy), and how many replicas exist (N, the replication factor). The next logical question is: when I perform a read or write, how many of those N replicas need to acknowledge the operation for it to be considered successful?
In the next lesson, we will dive into Cassandra's eventual consistency model, exploring the roles of tunable consistency levels for writes (W) and reads (R). This will complete the foundational triad of N, W, and R that governs Cassandra's trade-offs between consistency, availability, and performance.