Skip to main content
Create your own

Replicating Kafka Topics with MirrorMaker 2

Hello! Welcome to our next lesson.

In our previous session, we deconstructed the architecture of MirrorMaker 2, analyzing its role in geo-replication, the trade-offs of its asynchronous nature (defining your RPO), and the complexities of consumer failover. We established that MM2 is essentially a specialized Kafka Connect cluster dedicated to replicating data between independent Kafka clusters.

Today, we transition from theory to practice. Your goal is to configure and run MirrorMaker 2 to replicate a topic between two Kafka clusters. This hands-on lab will solidify your understanding of how MM2 operates and provide a tangible foundation for designing and managing multi-datacenter systems.

1. The MirrorMaker 2 Configuration File

The behavior of MirrorMaker 2 is controlled by a single properties file. Before we launch anything, it's crucial to understand the key directives that define the replication topology.

Let's examine a typical configuration for a unidirectional, active-passive replication flow.

Kafka MirrorMaker—Tutorial, best practices and alternatives

The article 'Kafka MirrorMaker—Tutorial, best practices and alternatives' from RedPanda provides a clear, annotated example of an MM2 configuration file. We will use this as the basis for our setup.

Please review the sample configuration file found in the section 'MirrorMaker setup'. Focus on understanding the purpose of the main configuration blocks.

Let's break down the most important properties from that example:

  • Cluster Aliases and Connection:

    clusters = A, B
    A.bootstrap.servers = localhost:9092
    B.bootstrap.servers = localhost:9093
    

    Here, A and B are logical aliases for our source and target clusters. You can name them anything, but descriptive names like us-east-1 and eu-west-1 are common in production. Each alias is then mapped to its respective Kafka brokers.

  • Replication Flow Definition:

    A->B.enabled = true
    A->B.topics = .*
    

    This is the core of the replication logic. A->B.enabled = true explicitly creates a replication flow from cluster A to cluster B. The topics property specifies which topics to replicate, using a regular expression. .* means "replicate all topics." In a real system, you would likely use a more specific regex, like payments-.* or ^(public|shared)\\..*, to control costs and data flow.

  • Internal Topic Configuration:

    checkpoints.topic.replication.factor=1
    heartbeats.topic.replication.factor=1
    offset-syncs.topic.replication.factor=1
    

    As we discussed, MM2 creates its own internal topics to function. These include heartbeats to check cluster connectivity and checkpoints to store consumer group offset mappings for failover. For this lab, a replication factor of 1 is sufficient, but in production, these would be set to 3 or more for fault tolerance, just like any other critical topic.

2. Practical Lab: Setting up Cross-Cluster Replication

Now, let's build and run this setup. This lab will guide you through creating two Kafka clusters in Docker and using MM2 to replicate data from one to the other.

Step 2.1: Create the Configuration File

First, create a file named mm2.properties on your local machine and paste the following content into it. This is based on the resource we just reviewed.

# mm2.properties

# Define cluster aliases
clusters = A, B

# Connection info for source cluster 'A'
A.bootstrap.servers = localhost:9092

# Connection info for target cluster 'B'
B.bootstrap.servers = localhost:9093

# Enable replication from A to B
A->B.enabled = true

# Replicate all topics from A to B
A->B.topics = .*

# For this lab, we'll disable the default topic prefixing.
# This makes consumer failover simpler to observe, but as we discussed,
# it's risky in active-active setups.
replication.policy.class = org.apache.kafka.connect.mirror.IdentityReplicationPolicy

# --- Internal Topic Settings for a Lab Environment ---
replication.factor=1
checkpoints.topic.replication.factor=1
heartbeats.topic.replication.factor=1
offset-syncs.topic.replication.factor=1
offset.storage.replication.factor=1
status.storage.replication.factor=1
config.storage.replication.factor=1

Note: I've added replication.policy.class = org.apache.kafka.connect.mirror.IdentityReplicationPolicy. This overrides the default behavior of prefixing replicated topics (e.g., A.my-topic). For an active-passive DR scenario, having identical topic names can simplify client failover logic, as applications don't need to be aware of which cluster they are pointing to. However, this policy is dangerous in an active-active setup as it can cause infinite replication loops.

Step 2.2: Launch the Kafka Clusters

We will now use Docker to launch two single-node Kafka clusters. The following commands are adapted from the "Kafka MirrorMaker setup tutorial" resource.

Open your terminal.

  1. Launch Zookeeper and Kafka for Cluster A (Source):

    # Zookeeper for Cluster A
    docker run -d --name zookeeper-a -p 2181:2181 confluentinc/cp-zookeeper:7.3.2
    
    # Kafka Broker for Cluster A
    docker run -d --name kafka-a -p 9092:9092 \
    -e KAFKA_BROKER_ID=1 \
    -e KAFKA_ZOOKEEPER_CONNECT=zookeeper-a:2181 \
    -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP="PLAINTEXT:PLAINTEXT,PLAINTEXT_HOST:PLAINTEXT" \
    -e KAFKA_ADVERTISED_LISTENERS="PLAINTEXT://localhost:9092,PLAINTEXT_HOST://kafka-a:29092" \
    -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \
    -e KAFKA_GROUP_INITIAL_REBALANCE_DELAY_MS=0 \
    -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \
    -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \
    --link zookeeper-a \
    confluentinc/cp-kafka:7.3.2
    
  2. Launch Zookeeper and Kafka for Cluster B (Target):
    Note the port changes (2182, 9093) to avoid conflicts. We also mount the directory containing your mm2.properties file into the container.

    # Zookeeper for Cluster B
    docker run -d --name zookeeper-b -p 2182:2181 confluentinc/cp-zookeeper:7.3.2
    
    # Kafka Broker for Cluster B
    docker run -d --name kafka-b -p 9093:9092 \
    -v $(pwd):/config \
    -e KAFKA_BROKER_ID=2 \
    -e KAFKA_ZOOKEEPER_CONNECT=zookeeper-b:2181 \
    -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP="PLAINTEXT:PLAINTEXT,PLAINTEXT_HOST:PLAINTEXT" \
    -e KAFKA_ADVERTISED_LISTENERS="PLAINTEXT://localhost:9093,PLAINTEXT_HOST://kafka-b:29092" \
    -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \
    -e KAFKA_GROUP_INITIAL_REBALANCE_DELAY_MS=0 \
    -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \
    -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \
    --link zookeeper-b \
    confluentinc/cp-kafka:7.3.2
    

    Replace $(pwd) with the absolute path to the directory containing mm2.properties if you are on Windows or if your shell requires it.

Give the containers a minute to start up. You can check their status with docker ps.

Step 2.3: Start MirrorMaker 2

We will run the MM2 process inside the kafka-b (target) container. This aligns with the best practice of deploying MM2 in the same datacenter as the target cluster to minimize latency on the produce side.

docker exec -it kafka-b connect-mirror-maker /config/mm2.properties

You will see a stream of logs as MM2 starts up, initializes its connectors, and connects to both clusters.

Step 2.4: Verify Replication

Now, let's test the pipeline. You will need two more terminal windows.

  1. In Terminal 1, create a topic in Cluster A:

    docker exec -it kafka-a kafka-topics --bootstrap-server localhost:9092 --create --topic orders --partitions 1 --replication-factor 1
    
  2. In Terminal 2, list topics in Cluster B:
    Wait a few seconds for MM2 to detect the new topic, then run:

    docker exec -it kafka-b kafka-topics --bootstrap-server localhost:9093 --list
    

    You should see the orders topic, along with several MM2 internal topics (e.g., heartbeats, mm2-offset-syncs...). This confirms MM2 has successfully replicated the topic creation.

  3. In Terminal 1, start a producer for the orders topic in Cluster A:

    docker exec -it kafka-a kafka-console-producer --bootstrap-server localhost:9092 --topic orders
    

    Type a few messages and press Enter after each one.

    >{"order_id": 101, "amount": 150.00}
    >{"order_id": 102, "amount": 75.50}
    
  4. In Terminal 2, start a consumer for the orders topic in Cluster B:

    docker exec -it kafka-b kafka-console-consumer --bootstrap-server localhost:9093 --topic orders --from-beginning
    

    You should see the messages you produced to Cluster A appear in the consumer attached to Cluster B. This demonstrates that the end-to-end replication is working.

You have now successfully configured and run a cross-cluster replication pipeline with MirrorMaker 2. Feel free to experiment by stopping the kafka-a container to simulate an outage and observing the logs in the MM2 process.

3. Deployment Best Practices

This lab provides a functional setup, but moving to a production environment requires additional considerations. Your experience with high-load systems highlights the importance of operational robustness.

Kafka MirrorMaker—Tutorial, best practices and alternatives

The RedPanda article concludes with a section on best practices that are essential for a production-grade MM2 deployment.

Please read the section 'Kafka MirrorMaker best practices'. Focus on the recommendations for deployment location and performance tuning.

The key takeaways reinforce our architectural principles:

  • Target-based Deployment: Running MM2 in the target datacenter is crucial. It ensures that if the WAN link between datacenters is flaky, messages are buffered closer to their final destination, making the write to the target cluster more reliable.
  • Performance Tuning & Monitoring: MM2 is a consumer of your source cluster. You must provision sufficient capacity (network bandwidth, broker CPU) on the source cluster to handle this additional read load without impacting your primary applications. Monitoring the replication lag (by tracking the MM2 consumer group's offsets on the source cluster) is the most critical metric for understanding your RPO at any given moment.

Conclusion

In this lesson, you moved from architectural theory to hands-on implementation. You configured and deployed a complete MirrorMaker 2 pipeline, replicating data between two distinct Kafka clusters.

Key Takeaways:

  • MM2 is configured via a properties file that defines cluster aliases, connection details, and replication flows (A->B).
  • The replication.policy.class property is critical for controlling topic naming strategy, with IdentityReplicationPolicy being useful for active-passive DR but dangerous for active-active setups.
  • The best practice is to run the MM2 process in the target datacenter to improve the reliability of writes to the target cluster.
  • Verifying replication involves creating a topic and producing messages on the source, then observing the topic and consuming messages on the target.

Next Up

We have successfully established a data pipeline between two clusters. However, in any distributed system, failures and retries are inevitable. This can lead to duplicate messages. How do we ensure that a message is processed exactly once, even in the face of network errors? In our next lesson, we will address this by learning how to implement an idempotent Kafka producer to achieve exactly-once semantics per partition.

Can't find a good explanation? Sign up and we'll make it for you

Sign up