Hello! Welcome to the final lesson in our module on Event-Driven Architecture.
In our previous session, we tackled the challenge of ensuring exactly-once processing by designing idempotent event handlers. This is a crucial piece of the puzzle for building robust systems. However, a system that cannot change is a system that cannot last. Today, we address the other side of that coin: evolution.
Introduction
Lesson 5: Outlining a Strategy for Non-Breaking Event Schema Evolution
Learning Outcome: By the end of this lesson, you will be able to outline a strategy for non-breaking event schema evolution.
In any long-lived, high-load system, especially in domains like payments or fintech where you have worked, business requirements change, bugs are fixed, and optimizations are made. These changes often necessitate evolving the "data contracts"—the schemas of the events that services exchange.
Doing this without causing downtime or data corruption is a significant architectural challenge. A poorly managed schema change can break downstream consumers, leading to outages and data integrity issues. This is analogous to managing breaking changes in a public API, but in an event-driven world, the "consumers" can be numerous, diverse, and not always under the direct control of a single team.
Our goal today is to formulate a clear, practical strategy for evolving event schemas gracefully, ensuring that your distributed system can adapt and grow without requiring "big bang" deployments or coordinated downtime.
1. What Makes a Schema Change "Breaking"?
Before we can devise a strategy to avoid breaking changes, we need a precise definition of what they are. The distinction is not always as simple as "adding vs. removing a field," as subtle semantic changes can be just as disruptive.
How to Handle Breaking Changes for API and Event ...
To establish a clear foundation, let's start with this article from freeCodeCamp that provides excellent definitions and examples of breaking versus non-breaking changes.
Please read the sections 'What are Breaking Changes?' and 'What are Non-Breaking Changes?'. As you read, focus on the distinction between structural changes (like renaming a field) and semantic changes (like changing a unit from dollars to cents).
As the article clarifies, a breaking change is any modification that forces a consumer to update its code. This includes:
- Structural Changes:
- Renaming an attribute (
orderId->order_id) - Removing an attribute
- Changing a data type (
string->integer) - Making an optional attribute required
- Renaming an attribute (
- Semantic Changes:
- Changing the format of data (date
mm/dd/yyyy->yyyy-mm-dd) - Changing the meaning or unit (
amountin dollars ->amountin cents) - Adding stricter validation constraints (e.g., a new max length on a string)
- Changing the format of data (date
Conversely, non-breaking changes are generally additive or permissive:
- Adding a new, optional attribute.
- Making a required attribute optional.
The key insight is that consumers are often programmed to be "liberal in what they accept." They can typically ignore new fields they don't recognize, but they will fail if a field they depend on is removed, renamed, or has its type changed unexpectedly.
2. Formalizing Evolution: Schema Compatibility Models
In a large-scale system using a message broker like Kafka, managing schema evolution manually is error-prone. This is why a Schema Registry is a critical piece of infrastructure. It acts as a centralized repository for your event schemas and, more importantly, can be configured to enforce compatibility rules automatically when a new schema version is published.
To use a schema registry effectively, you must understand the formal compatibility models it can enforce.
Best Practices for Evolving Schemas in Schema Registry
This documentation from Solace provides a clear and concise explanation of the standard compatibility types used in schema registries. The concepts are directly applicable to the popular Confluent Schema Registry used with Kafka.
Read the sections 'Key Concepts of Dynamic Schema Evolution' and 'Compatibility Rules and Strategies'. Pay close attention to the definitions of Backward, Forward, and Full compatibility. The comparison table is particularly useful for summarizing the rules.
Let's summarize these crucial concepts, as they dictate your deployment strategy:
-
Backward Compatibility: A new schema version is backward compatible if consumers using the new schema can correctly process data produced with an older schema.
- Rule: You can remove fields or add new optional fields (or fields with default values).
- Impact: Allows you to upgrade consumers before producers. This is a very common and practical strategy.
-
Forward Compatibility: A new schema version is forward compatible if consumers using an older schema can correctly process data produced with the new schema.
- Rule: You can add new fields or remove optional fields. (Consumers must ignore unknown fields).
- Impact: Allows you to upgrade producers before consumers. This is useful when you have many consumers that are difficult to update simultaneously.
-
Full Compatibility: The schema is both backward and forward compatible.
- Rule: You can only add or remove optional fields.
- Impact: Allows producers and consumers to be upgraded in any order. This offers the most flexibility but is also the most restrictive.
For many organizations, Backward compatibility is the chosen default. It creates a clear process: update all consumers to handle the new schema (often by providing default values for new fields), and only then update producers to start sending it.
3. A Strategy for Non-Breaking Schema Evolution
Now, let's combine these concepts into a practical, step-by-step strategy.
Step 1: Use a Schema Registry and Choose a Compatibility Mode
The first and most important step is to not leave this to chance.
- Adopt a Schema Registry (like Confluent Schema Registry for Kafka).
- Configure a default compatibility mode for your topics.
BACKWARDis a safe and common starting point. This provides an automated guardrail against accidental breaking changes.
Step 2: Adopt Additive and Permissive Design Principles
Your development teams should operate under a set of principles that align with your chosen compatibility mode.
Best Practices for Evolving Schemas in Schema Registry
Let's return to the Solace document to see how these principles are implemented in practice using schema languages like Avro or JSON Schema.
Please read the sections 'Schema Design Principles for Evolution', the implementation examples under 'Optional Fields with Default Values', and finally 'Avoiding Anti-Patterns'. This will provide concrete techniques for your schemas.
The core principles to enforce are:
- Use Default Values: When adding a new field that is essential for new consumers, providing a sensible default value is the key to backward compatibility. New consumers can process old events by substituting the default.
- Make New Fields Optional: If a new field is not strictly required, define it as optional. In Avro or JSON Schema, this is often done with a union type, e.g.,
["null", "string"]. - Avoid Renaming: Renaming a field is a breaking change. The correct approach is to add the new field and temporarily support both.
- Use Semantic Versioning: Version your schemas (e.g.,
order-created-v1.1.avsc). Use major version bumps (v2.0) to signal deliberate, coordinated breaking changes, while minor/patch versions (v1.1,v1.0.1) must remain compatible.
Step 3: Implement the Expand-Contract Pattern for Major Changes
What if you absolutely must rename or remove a field, or make a significant semantic change? You can't do it in one step. The solution is a two-phase pattern, which you may recognize from database migrations.
Let's use the example of renaming a field from customerId to userId.
Phase 1: Expand
- Schema Change: Add the new field
userIdto the event schema as an optional field. This change is backward compatible. - Producer Deployment: Deploy new producer code that writes to both
customerId(for old consumers) anduserId(for new consumers). - Consumer Deployment: Deploy new consumer code. This code must be written to be resilient: it should prefer reading from
userId, but if that field is null or absent (indicating an older event), it must fall back to reading fromcustomerId. - (Optional) Backfill: If necessary for analytics or reprocessing, run a job that reads old events and republishes them with the
userIdfield populated.
At the end of this phase, all running code can handle both old and new event formats.
Phase 2: Contract
- Producer Deployment: Once you have verified that all consumers have been updated to the code from Phase 1, deploy a new version of the producer that writes only to the
userIdfield. ThecustomerIdis no longer populated. - Consumer Deployment: Deploy a final version of the consumer code that now reads only from
userId. The fallback logic can be removed, simplifying the code. - Schema Change: Finally, you can publish a new schema version that formally removes the
customerIdfield.
This methodical, multi-step process ensures that at no point is a consumer unable to process an event, achieving a zero-downtime schema migration.
Conclusion
Managing schema evolution is a process that requires discipline, tooling, and a clear strategy. Simply hoping that developers won't make breaking changes is not a strategy. By formalizing the process, you enable your event-driven architecture to adapt over time without sacrificing stability.
Key Takeaways:
- Non-breaking changes are additive and permissive. The most common non-breaking change is adding a new, optional field.
- A Schema Registry is a critical tool for automating the enforcement of compatibility rules.
- Compatibility models (Backward, Forward, Full) provide a formal framework for managing evolution and dictate the order in which you can deploy producers and consumers.
BACKWARDis a common and practical choice. - A robust strategy combines:
- Using a schema registry with a default compatibility mode.
- Following design principles like using default values and optional fields.
- Applying the Expand-Contract pattern for complex changes like renames or removals.
This lesson concludes our module on Event-Driven Architecture. We have journeyed from modeling business processes with events, to building read-model projections, ensuring robust processing with idempotent handlers, and now, managing the inevitable evolution of event schemas. These concepts are the pillars of building scalable, resilient, and maintainable event-driven systems.