Skip to main content
Create your own

Kafka Acknowledgment Levels and Trade-offs

Hello! Welcome back to our series on designing high-load distributed systems.

In our previous lesson, we explored Kafka's server-side replication architecture, focusing on the In-Sync Replica (ISR) set and the High Watermark (HWM). We established that these mechanisms are the foundation for Kafka's durability and consistency guarantees, ensuring that a message is only considered "committed" and visible to consumers after it's safely replicated.

Today, we shift our focus from the server to the client. While the broker manages replication, the producer has direct control over the durability guarantees it requires for each message it sends. This is configured through a single, powerful parameter: acks.

Lesson Goal: By the end of this session, you will be able to configure producer acknowledgment levels (acks=0, 1, all) and analyze the critical trade-offs between write latency and data durability for each setting.

1. The Producer's Durability Contract: The acks Configuration

The acks (acknowledgments) configuration determines how many broker acknowledgments the producer must receive before it considers a write request successful. This setting is the primary control you have as a developer to balance the speed of your writes against the risk of losing data.

Let's examine the three possible settings and their implications.

This diagram provides a high-level overview of the three acknowledgment modes. It illustrates the message flow and highlights the fundamental trade-off: as durability increases from `acks=0` to `acks=all`, so does the latency of the produce request.

We will now break down each of these modes in detail.

2. acks=0: Fire and Forget

This setting provides the lowest latency but also the weakest durability guarantee.

  • Behavior: The producer sends the message to the broker and does not wait for any acknowledgment. The send call returns as soon as the message is written to the producer's underlying network buffer.
  • Latency: Extremely low. The producer experiences no network round-trip delay.
  • Durability: None. There is no guarantee the message will reach the broker. It could be lost due to a transient network issue, an ongoing leader election, or if the broker crashes before persisting the message. The producer will not be notified of the failure.
  • Use Cases: Situations where occasional data loss is acceptable in exchange for maximum throughput and minimal latency. Examples include collecting high-volume, low-value telemetry, metrics, or non-critical logs.

3. acks=1: Leader Acknowledgment

This is the default setting in many older Kafka client versions and offers a middle ground between performance and durability.

  • Behavior: The producer sends the message and waits for an acknowledgment from the partition leader only. The leader writes the message to its own log file and then sends the acknowledgment. It does not wait for the follower replicas to copy the data.
  • Latency: Moderate. It includes the network round-trip time for the message to travel to the leader and for the acknowledgment to return.
  • Durability: Good, but not perfect. Acknowledgment from the leader confirms the message has reached the cluster. However, a small window for data loss exists: if the leader broker fails after sending the acknowledgment but before the followers in the ISR have replicated the message, the message will be lost. The producer will think the write was successful, but the new leader (elected from the remaining ISR members) will not have the data.
This diagram illustrates the `acks=1` flow. The producer sends messages to the leader (Broker-101). The leader writes them to its local log and immediately acknowledges the producer. The replication to followers (Broker-102, Broker-103) happens asynchronously in the background.

4. acks=all (or -1): Full ISR Acknowledgment

This setting provides the strongest durability guarantee available in Kafka.

  • Behavior: The producer sends the message and waits for the leader to acknowledge it. The leader will only send the acknowledgment after it has written the message to its own log and has received confirmation from all followers in the current In-Sync Replica (ISR) set that they have also written the message to their logs.
  • Latency: Highest. The total time includes the round-trip to the leader, plus the time for the data to be replicated to all ISR followers and for their acknowledgments to return to the leader.
  • Durability: Highest. This setting ensures that as long as at least one replica from the ISR remains available, the committed message will not be lost. If the leader fails, any other replica from the ISR can be promoted to the new leader, and it is guaranteed to have the message.
  • Use Cases: Systems where data loss is not an option. This is the standard choice for financial systems (like the payment and trading systems you've built), critical business event streams, and any application using event sourcing.

To solidify your understanding of these modes, please review the following resources.

Kafka Acknowledgment Settings Explained: acks=0,1,all

The article 'Kafka Acknowledgment Settings Explained' from Dattell provides a concise and clear breakdown of each acks setting, including its behavior, pros, cons, and ideal use cases.

Please read the sections for acks=0, acks=1, acks=all, and the summary section 'Choosing the Right acks Setting'. This will give you a structured overview of the trade-offs.

5. Enforcing Durability: acks=all and min.insync.replicas

Using acks=all is a request from the producer for maximum durability. However, to create a truly robust system, you must enforce this guarantee on the server side as well. What if the ISR shrinks to just the leader due to network issues or slow followers? In that case, acks=all becomes functionally equivalent to acks=1, as there are no followers to wait for. This silently degrades your durability guarantee.

To prevent this, you must combine acks=all with the topic-level configuration min.insync.replicas.

This setting specifies the minimum number of replicas that must be in the ISR for the partition to accept writes when the producer uses acks=all.

Let's analyze the interaction with an example:

  • Topic configuration: replication.factor = 3, min.insync.replicas = 2
  • Producer configuration: acks = all

Scenario 1: Healthy Cluster
The ISR contains 3 replicas (the leader and 2 followers). A produce request arrives. The leader replicates to the 2 followers. Once all 3 have the message, the leader sends an ack. The write succeeds because the ISR size (3) is >= min.insync.replicas (2).

Scenario 2: One Follower Fails
One follower broker goes down. The ISR shrinks to 2 replicas (the leader and 1 follower). A produce request arrives. The leader replicates to the remaining follower. Once both have the message, the leader sends an ack. The write succeeds because the ISR size (2) is still >= min.insync.replicas (2). The system remains available and durable.

Scenario 3: Two Followers Fail
Two follower brokers go down. The ISR shrinks to just 1 replica (the leader). A produce request arrives. Now, the broker rejects the write and the producer receives a NotEnoughReplicas or NotEnoughReplicasAfterAppend exception. The write fails because the ISR size (1) is < min.insync.replicas (2).

This is a critical design choice: the system chooses to become unavailable for writes rather than risk violating its durability contract. For the systems you've worked on, this is the expected behavior for handling critical financial data.

The following reading explains this powerful combination.

How to Tune Kafka's Durability and Ordering Guarantees

The Confluent documentation 'How to Tune Kafka's Durability' explains how acks and min.insync.replicas work together to enforce durability guarantees.

Please read the sections 'Producer acks = 0', 'Producer acks = 1', 'Producer acks = all', and 'Topic min.insync.replicas'. Focus on how the min.insync.replicas setting provides a server-side enforcement of the durability requested by acks=all.

Summary of Trade-offs

acks Setting Latency Throughput Durability Guarantee Typical Use Case
acks=0 Lowest Highest None (at-most-once) Metrics, logging, telemetry
acks=1 Medium High Leader-only (risk on leader failover) General-purpose, where minimal risk is ok
acks=all Highest Lower Full ISR (at-least-once) Financial transactions, critical events

Conclusion

In this lesson, we've analyzed the producer's acks configuration, the primary mechanism for controlling the trade-off between write latency and durability in Kafka.

Key Takeaways:

  • The acks setting determines how many brokers must acknowledge a write before the producer considers it successful.
  • acks=0 ("fire and forget") offers the highest performance with no durability guarantees.
  • acks=1 (leader ack) offers a balance of good performance with a small, but present, risk of data loss on leader failure.
  • acks=all (ISR ack) provides the strongest durability guarantee by ensuring the message is replicated to all in-sync replicas before acknowledgment.
  • For mission-critical applications, acks=all must be paired with min.insync.replicas to enforce the durability guarantee on the server side, preferring unavailability over potential data loss.

Next Up

Using acks=all ensures your messages are not lost (at-least-once delivery). However, network issues or client-side retries can cause the same message to be written to Kafka more than once, leading to duplicate processing. In our next lesson, we will address this by exploring how to implement an idempotent Kafka producer, the first step towards achieving exactly-once semantics.

Can't find a good explanation? Sign up and we'll make it for you

Sign up