Hello! Welcome to our next lesson on Kafka.
In our last session, we explored how to achieve scalable and resilient consumption using consumer groups and partition assignment strategies. We saw that the StickyAssignor combined with static group membership provides the most robust solution for modern high-load systems by minimizing rebalancing disruptions.
Now that we understand how data is produced and consumed, a critical operational question remains: what happens to the data stored in Kafka's logs over time? Unmanaged, this data would grow indefinitely, consuming all available disk space. This brings us to today's topic: managing the data lifecycle within Kafka.
Lesson Goal: Today, you will learn how to configure log retention policies: time-based and size-based. We will cover the different retention strategies, the hierarchy of configuration, and the specific parameters you'll use to control disk usage, meet compliance requirements, and manage costs effectively in your Kafka clusters.
This is a fundamental aspect of operating Kafka in production, directly aligning with your goal of mastering the practical configuration of high-load systems.
1. The Role of Log Retention
Before diving into configuration, it's important to understand why retention is so critical. Kafka's power comes from its durable, replayable log. However, this durability comes at the cost of storage. A well-defined retention strategy is essential for:
- Controlling Storage Costs: Preventing unbounded disk growth is the most immediate operational concern.
- Ensuring Data Availability: Retention periods must be long enough for all consumers, including slow or offline ones, to process the data.
- Meeting Compliance Requirements: Many industries (like finance) have legal mandates to retain data for specific periods (e.g., 7 years for transaction records).
- System Performance: While Kafka is designed for large logs, excessive data can still impact performance, for example, by increasing recovery times after a broker restart.
To begin, let's get a formal overview of Kafka's retention policies and the underlying cleanup mechanism.
Apache Kafka® Retention Explained: Policies & Best ...
The article 'Apache Kafka® Retention Explained' from Confluent provides an excellent introduction to the concept of retention and the two primary cleanup policies: deletion and compaction.
Please read the introductory sections 'Why Apache Kafka® Retention Matters' and 'Kafka Retention Policy'. Then, jump to the section 'How Kafka Handles Log Cleanup' and read the subsections 'Why Log Cleanup Matters' and 'Log Deletion: Retention-Based Cleanup'. For now, you can skim the part about 'Log Compaction'; we are focusing on the default 'delete' policy.
As you read, note that today's lesson focuses entirely on the cleanup.policy=delete setting, which is the default for all Kafka topics. This policy discards old data based on time or size, which we will now learn to configure.
2. Configuring Retention Policies
Kafka provides a flexible, hierarchical system for configuring retention. You can set global defaults for the entire cluster and then override them for specific topics that have different requirements.
Configuration Hierarchy
The settings are applied with the following precedence, from highest to lowest:
- Topic-level configuration: Explicitly set when creating or altering a topic. This always wins.
- Broker-level configuration: Defaults set in the
server.propertiesfile for each broker. These apply to any topic that doesn't have its own explicit setting. - Kafka default values: Hardcoded defaults in the Kafka software, used if a setting is specified at neither the topic nor the broker level.
This hierarchy allows you to set a sensible default for most topics (e.g., retain data for 7 days) while accommodating special cases (e.g., a temporary topic for a data migration might only need 24 hours of retention).
The following reading details the specific parameters and how to apply them at different levels.
Apache Kafka® Retention Explained: Policies & Best ...
This next section of the Confluent article provides a detailed guide to the configuration parameters for both time-based and size-based retention.
Please read the entire section 'How to Configure Kafka Retention'. Pay close attention to: The 'Key Time-Based Parameters' (log.retention.ms, log.retention.hours) and their precedence. The 'Essential Size-Based Parameters' (log.retention.bytes). The examples for Broker-Level and Topic-Level configuration. The Java Admin API example, which may be relevant given your background.
Key Parameters Summary
Let's consolidate the key parameters you've just read about:
- Time-Based Retention:
retention.ms: The maximum time to retain a log segment, in milliseconds. This has the highest precedence.retention.minutes: Retention in minutes.retention.hours: Retention in hours. (Default: 168, i.e., 7 days).
- Size-Based Retention:
retention.bytes: The maximum total size of all log segments for a partition before old segments are deleted. (Default: -1, meaning no size limit).
- Cleanup Check:
log.retention.check.interval.ms: How frequently the log cleaner checks if any segments can be deleted. (Default: 300000, i.e., 5 minutes).
Important Interaction: When both retention.ms and retention.bytes are set for a topic, Kafka's log cleaner will delete segments if either condition is met. For example, if you set retention to 7 days or 100 GB, data will be deleted if it's older than 7 days or if the partition size exceeds 100 GB.
3. Practical Configuration and Verification
Now let's look at the concrete commands for applying and inspecting these settings.
Broker-Level Configuration
To set a cluster-wide default, you edit the server.properties file on each Kafka broker and then perform a rolling restart. For instance, to set a default retention of 30 days and a size limit of 500 GB per partition:
# server.properties
log.retention.hours=720
log.retention.bytes=536870912000
Topic-Level Configuration
More commonly, you will manage retention on a per-topic basis using the kafka-topics.sh command-line tool.
The following reading provides clear examples of creating a topic with custom retention and inspecting its configuration.
Kafka Logging Guide: Advanced Concepts
The 'Kafka Logging Guide' from CrowdStrike offers concise, practical examples of applying retention policies at the topic level.
Please read the section 'Creating and Inspecting a Compacted Topic'. Although the title mentions 'compacted', the examples demonstrate setting retention.ms and retention.bytes, which are the focus of our lesson. Pay attention to the --config flag for creation and the output of the --describe command.
Let's walk through some common scenarios using the commands you've just seen.
Scenario 1: Financial Transaction Logs
- Requirement: Retain all transaction data for 7 years for compliance. Size is not a primary concern, but time is absolute.
- Command:
Here, we leave# 7 years in milliseconds = 7 * 365.25 * 24 * 60 * 60 * 1000 ≈ 220903200000 kafka-topics.sh --create --topic financial-transactions \ --bootstrap-server localhost:9092 \ --partitions 50 \ --replication-factor 3 \ --config retention.ms=220903200000retention.bytesat its default of -1 (unlimited).
Scenario 2: Real-time Analytics Events
- Requirement: Keep only the most recent 10 GB of data for a dashboard. The age of the data is less important than its total volume.
- Command:
# 10 GB in bytes = 10 * 1024 * 1024 * 1024 = 10737418240 kafka-topics.sh --create --topic dashboard-metrics \ --bootstrap-server localhost:9092 \ --partitions 12 \ --replication-factor 3 \ --config retention.bytes=10737418240
Scenario 3: Modifying an Existing Topic
- Requirement: The
dashboard-metricstopic is growing too fast. The business decides to reduce its retention to 24 hours, in addition to the size limit. - Command:
Now, this topic will delete data if it's older than 24 hours OR if the partition size exceeds 10 GB.# 24 hours in milliseconds = 86400000 kafka-topics.sh --alter --topic dashboard-metrics \ --bootstrap-server localhost:9092 \ --config retention.ms=86400000
Verification:
You can always verify the current configuration of a topic:
kafka-topics.sh --describe --topic dashboard-metrics \
--bootstrap-server localhost:9092
The output will include a Configs section showing the applied retention policies.
Conclusion
In this lesson, we've covered the critical operational task of managing Kafka's data lifecycle through retention policies. You are now equipped to make informed decisions about how long to store data in your topics, balancing performance, cost, and business requirements.
Key Takeaways:
- Log retention is a fundamental mechanism for managing disk usage and cost in a Kafka cluster.
- The default cleanup policy is
delete, which removes data based on time or size. - Time-based retention is configured with
retention.ms(or.hours/.minutes). - Size-based retention is configured with
retention.bytes. - When both are set, data is deleted when the first limit is reached.
- Configurations can be set globally at the broker level or overridden for specific use cases at the topic level.
Next Up
We have now established a solid understanding of how to manage a single Kafka cluster. However, high-load, globally distributed systems often require multi-cluster deployments for disaster recovery, geographic locality, or regulatory compliance. In our next lesson, we will begin to explore this by analyzing the architecture and trade-offs of cross-datacenter replication with MirrorMaker 2, Kafka's tool for replicating data between clusters.