Skip to main content
Create your own
Lesson illustration

Message Queues vs. Distributed Logs

Welcome back! In our last lesson, we established the fundamental divide between synchronous and asynchronous communication, concluding that asynchronous patterns are essential for building resilient and scalable systems. We introduced message brokers as the critical intermediary that decouples services and absorbs shocks to the system.

Now, we'll dive deeper into the technologies that power these asynchronous architectures. This lesson addresses a question that frequently appears in system design interviews and is crucial for any architect: what is the difference between a message queue like RabbitMQ and a distributed log like Kafka? While both can be used to send messages, they are built on entirely different philosophies. Understanding this distinction is key to making informed architectural decisions that will support your system's growth for years to come.

1. Two Philosophies: The Post Office vs. The Immutable Log

At the heart of the Kafka vs. RabbitMQ discussion lies a fundamental difference in their conceptual models. Mistaking one for the other is a common pitfall that can lead to poorly designed systems.

  • RabbitMQ is a message broker. Think of it as a smart post office. Its primary job is to receive messages (letters) from producers, use a complex set of routing rules (the postal system) to deliver them to the correct consumers (mailboxes), and ensure they are received. Once a message is successfully delivered and acknowledged, the broker's job is done, and the message is typically deleted. It's designed for work distribution.

  • Apache Kafka is a distributed streaming platform. It's better to think of it as an immutable, append-only log or a library's historical archive. Its primary job is to record events. Producers append events to the end of a log (the archive), where they persist for a configured amount of time. Consumers read from this log at their own pace, keeping track of their own position (a bookmark). The event is not deleted when it's read. This design is optimized for event streaming and replayability.

This illustration captures the core architectural difference. RabbitMQ focuses on routing messages to a destination queue, where they are consumed. Kafka organizes messages into partitioned logs (topics) that multiple consumers can read from independently.

To solidify this core distinction, let's watch a short video that explains these two models from first principles.

System design from First Principle[11/15] : Kafka vs RabbitMQ vs SQS — Async System

The video "System design from First Principle" offers an excellent analogy that clarifies the difference between task queues (like RabbitMQ) and message streams (like Kafka).

Please watch the segment from the two main tools. As you watch, focus on the "post office" vs. "immutable scroll" analogy. Pay close attention to two key ideas: How messages are treated after being read (deleted vs. persisted). The implications for the number and type of consumers that can process the messages.

This fundamental difference—transient work item vs. durable event record—drives almost every other distinction between the two systems.

2. A Tale of Two Architectures

Now that we have the conceptual model, let's formalize it by looking at the key components of each architecture.

Apache Kafka vs. RabbitMQ: Comparing architectures ...

The article from quix.io provides a clear breakdown of the architectures of both systems, complete with diagrams.

First, read the section on Apache Kafka architecture, starting from this paragraph down to the end of the sub-section. Note the roles of topics, partitions, and brokers. The concept of partitions is key to Kafka's scalability. Next, read the section on RabbitMQ architecture, from the heading down to the end of the note on streams. Focus on understanding the flow: Publisher → Exchange → Binding → Queue → Consumer. The exchange is the "smart" part of the broker that handles complex routing logic.

Let's quickly summarize the key architectural components:

RabbitMQ: The Smart Broker

  • Exchanges: Receive messages from producers and decide where they should go. They come in several types (Direct, Topic, Fanout, Headers), providing powerful and flexible routing capabilities. This is where most of the "intelligence" of the broker lives.
  • Queues: The final destination for messages before they are delivered to consumers. They store messages until a consumer is ready.
  • Bindings: The rules that link an exchange to a queue.

This architecture follows a "smart broker, simple consumer" model. You configure the complex routing and delivery logic on the broker, and the consumer's job is relatively simple: connect to a queue and process the messages it receives.

Kafka: The "Dumb" Broker

  • Topics: A category or feed name to which events are published. You can think of a topic as a folder in a filesystem.
  • Partitions: Topics are split into multiple partitions. Each partition is an ordered, immutable sequence of records—a log. Partitions are the unit of parallelism in Kafka; they allow a topic to be scaled horizontally across multiple servers.
  • Offsets: A pointer that each consumer group maintains for each partition, marking how far it has read in the log.

This architecture promotes a "simple broker, smart consumer" model. The broker's main job is to efficiently write events to a log file and serve them to consumers. The consumer is responsible for tracking its own progress (offset), and therefore holds more of the logic.

3. Comparing Core Characteristics

The architectural differences we've just discussed lead to significant trade-offs in performance, message handling, and capabilities.

Kafka vs RabbitMQ

The video "Kafka vs RabbitMQ" from the "Hello Interview" channel provides a fantastic, concise summary of these trade-offs.

Please watch the video from the beginning of the Kafka explanation. This segment will walk you through the key differences: The "simple broker, smart consumer" model of Kafka and the concept of durable, replayable messages. The difference in ordering guarantees (global vs. per-partition). The implications for throughput and latency. The nuances of delivery guarantees (at-least-once vs. exactly-once).

Let's synthesize the information from the video and our other resources into a comparative table.

Characteristic RabbitMQ (Message Broker) Apache Kafka (Distributed Log)
Primary Model Smart Broker: Routes, tracks, and delivers messages which are then deleted. Optimized for work distribution. Dumb Broker: Appends events to a persistent, time-ordered log. Consumers track their own position. Optimized for event streaming.
Message Consumption Push-based (typically): Broker pushes messages to consumers. Once acknowledged (ACK), the message is deleted from the queue. Pull-based: Consumers pull messages in batches from the broker. Messages are retained based on policy and are not deleted on read.
Message Replay Not a native feature. Possible only for unacknowledged messages via a dead-letter queue. First-class feature. A consumer can "rewind" its offset to any point in the retention window and re-read historical data.
Throughput High, but can be limited by per-message overhead. Typically in the tens of thousands of messages/sec per broker. Extremely high, often millions of messages/sec. Optimized for sequential disk I/O and batching.
Latency Very low for moderate workloads, as the broker actively pushes messages. Slightly higher baseline latency due to the pull/batch model, but remains consistently low even under massive load.
Ordering Guaranteed message order within a single queue. With multiple consumers, global order is lost in favor of parallelism. Guaranteed order within a partition. There is no global order across partitions of a topic. This is the key trade-off for horizontal scalability.
Routing Complex & Flexible. Uses exchanges and bindings to support sophisticated routing patterns (e.g., content-based routing). Simple. Producers publish to topics. A partition key can be used to route all messages for a given entity to the same partition.
This table puts Kafka and RabbitMQ in the context of other messaging technologies like Amazon SQS and SNS, highlighting their distinct characteristics and best-fit use cases.

4. When to Use Which?

The final and most important question is: which tool should you choose for your project? The answer depends entirely on the problem you are solving.

Choose RabbitMQ when:

  • You need a task queue for background jobs (e.g., sending confirmation emails, processing image uploads, running reports). The "do this work and then forget it" model fits perfectly.
  • You require complex routing logic. For example, routing messages based on their content or headers to different consumers.
  • Your application requires low-latency message delivery for a moderate volume of messages.
  • Operational simplicity is a priority for a smaller-scale deployment.

Choose Apache Kafka when:

  • You are building an event-driven architecture where multiple, independent services need to react to the same stream of events (e.g., an OrderPlaced event consumed by shipping, billing, and analytics services).
  • You need to process high-throughput data streams in real-time (e.g., logging, metrics collection, IoT sensor data).
  • Message replayability is a requirement. For example, rebuilding a system's state from a history of events, or re-running an analytics job that had a bug.
  • You need a durable, long-term store of events that acts as a system of record.

It's also important to note that these tools are not mutually exclusive. A common pattern in large-scale systems is to use both: Kafka as the central event backbone (the durable log of everything that happens), with consumers that process these events and then place specific, actionable tasks onto RabbitMQ queues for worker processes.

Conclusion

In this lesson, we've moved beyond the general concept of asynchronous communication to dissect two of its most powerful implementations. You've learned that despite their superficial similarities, RabbitMQ and Kafka operate on fundamentally different principles.

Key Takeaways:

  • RabbitMQ is a message broker that excels at work distribution and complex routing. It acts like a post office, delivering a message to its destination and then removing it.
  • Kafka is a distributed streaming platform that excels at handling high-throughput event streams. It acts like an immutable log, recording events that can be read and re-read by multiple independent consumers.
  • The choice between them is a critical architectural decision driven by your use case: do you need to manage discrete tasks (RabbitMQ) or process a durable stream of historical events (Kafka)?
  • RabbitMQ follows a "smart broker, simple consumer" model, while Kafka uses a "simple broker, smart consumer" model. This dictates where the system's complexity lies.

In our next lesson, we will put this theory into practice. We'll start with the "post office" model and implement a producer-consumer pattern using RabbitMQ, giving you hands-on experience with a classic message queue.

Can't find a good explanation? Sign up and we'll make it for you

Sign up