Hello! Welcome back to our series on designing high-load distributed systems.
In our last lesson, we contrasted the choreography and orchestration approaches to implementing the Saga pattern. We established that choreographed Sagas rely on a decentralized, event-driven model where services communicate through a message broker, promoting loose coupling but making the overall workflow harder to trace.
A critical piece of the Saga pattern is its ability to maintain consistency in the face of failure. This is achieved through compensating transactions. Today, we will focus on the practical implementation of this rollback mechanism.
Our learning outcome is to implement compensation logic for a failed step in a choreographed Saga. We will explore:
- The event-driven mechanism that triggers compensation.
- How to design and implement compensation logic for different types of local transactions (database updates, API calls, etc.).
- The critical challenges and best practices, such as idempotency and reliability, that you must consider in a real-world system.
1. The Compensation Flow in a Choreographed Saga
In a choreographed Saga, there is no central controller to command a rollback. Instead, compensation is also a choreographed, event-driven process. When a service fails to complete its local transaction, it publishes a failure event. Upstream services, which have already completed their part of the Saga, must subscribe to these failure events to trigger their own compensating actions.
This diagram illustrates the concept. The green arrows (T1, T2, T3) represent the forward "happy path" transactions. If a step fails (e.g., the Payment service), it initiates a backward flow of compensating transactions (C1, C2, C3) by publishing a failure event.

Let's ground this in a more formal architectural view.
Saga choreography pattern - AWS Prescriptive Guidance
The AWS Prescriptive Guidance documentation provides a clear, high-level view of how this event flow works, defining the roles of both transactional and compensatory steps.
Please read the 'Implementation' section, focusing on the 'High-level architecture' subsection. Pay close attention to the sequence: the failure of the payment transaction (T3) leads to a compensatory transaction (C1), which in turn publishes a Payment failed message that triggers further compensations.
As you read, the key takeaway is that each service must not only know how to perform its part of the transaction but also how to undo it. The Inventory service is responsible for Inventory updated (T2) and Inventory reverted (C2). The trigger for C2 is an event published by a downstream service, like Payment failed.
This next diagram shows a concrete implementation of this pattern using AWS services.

2. Implementing Compensation Logic
Now for the core question: what does the code for a compensating transaction look like? The logic depends entirely on the actions performed by the corresponding local transaction. A local transaction isn't always just a database update; it can involve API calls, publishing messages, or a mix of operations.
The following resource provides excellent, practical code examples for each of these scenarios.
Implementing the Saga Pattern with Choreography and ...
Let's examine how to implement compensation logic in practice. This article breaks down the implementation based on the type of operation being compensated, with clear Java-based code snippets.
First, read the 'Rollback Mechanism' under the Choreography section to see the event subscriptions. Then, jump to the section 'Handling Local Transactions Beyond Database Interactions' and read subsections 1, 2, and 3. As you read, analyze the compensate...() methods and how they relate to the actions in the try blocks.
Let's synthesize what you've just read into a set of actionable patterns. A service that receives a failure event (e.g., PaymentFailed) must execute a compensating transaction that reverses its original action.
Pattern 1: Reverting a Database Operation
This is the most common case. The compensation logic is the inverse of the original database operation.
- If
T_iwasINSERT INTO orders ..., thenC_iisDELETE FROM orders WHERE .... - If
T_iwasUPDATE products SET stock = stock - 1, thenC_iisUPDATE products SET stock = stock + 1. - If
T_iinvolved a more complex state change, you might need to have stored the "before" state of the entity to restore it during compensation.
Pattern 2: Reverting an External API Call
As seen in the PaymentService example, if the local transaction involved calling an external, third-party API, the compensation must call another endpoint on that API to reverse the action.
- Transaction:
paymentGatewayClient.process(request) - Compensation:
paymentGatewayClient.refund(request)
This requires the external service to provide such a compensating API. If it doesn't, the Saga pattern may not be feasible for that interaction, or it may require manual intervention.
Pattern 3: "Reverting" a Published Message
You cannot retract a message that has already been published to a broker like Kafka. The compensation here is to publish a new event that semantically reverses the original one.
- Transaction: Publish
OrderCreatedevent. - Compensation: Publish
OrderCancelledevent.
Downstream services that acted on OrderCreated must now subscribe to OrderCancelled and run their own compensations.
Pattern 4: Handling Mixed Operations
Real-world services often perform multiple actions. The InventoryService example from the article is a perfect illustration. Its compensation logic must reverse every step of its local transaction.
- Transaction:
- Update local database:
inventoryRepository.updateStock(...) - Call external API:
warehouseClient.updateStock(...) - Publish event:
kafkaTemplate.send("InventoryUpdated", ...)
- Update local database:
- Compensation:
- Revert database update:
inventoryRepository.revertStock(...) - Call external API to undo change:
warehouseClient.revertStock(...) - Publish failure event:
kafkaTemplate.send("InventoryUpdateFailed", ...)
- Revert database update:
Notice that the atomicity of the original transaction (DB update + message publish) is critical, which is where the Transactional Outbox pattern we previously discussed becomes essential. The same applies to the compensating transaction.
3. Critical Challenges in Implementation
Implementing compensation logic introduces its own set of complexities that are vital to address in a production system. Your experience building low-latency and payment systems has likely exposed you to the importance of handling edge cases and failures gracefully.
Let's consult our resources again for a summary of these challenges.
Implementing the Saga Pattern with Choreography and ...
The article we just read concludes with a list of key considerations. The AWS documentation also provides a thorough list of issues. Let's review these to build a robust mental model.
Please read the 'Key Considerations' section of the Medium article.
Saga choreography pattern - AWS Prescriptive Guidance
Now, for a more architectural perspective on challenges, let's look at the AWS guide.
Read the 'Issues and considerations' section. Focus on the points regarding Idempotency, Transaction isolation, and Observability.
Based on these readings and your background, let's highlight the most pressing issues:
-
Idempotency of Compensations: A compensating action might fail due to a transient network error and be retried. The compensation logic must be idempotent. For example, a
refund()API call should be safe to call multiple times for the same transaction ID. Similarly,UPDATE products SET stock = stock + 1is not idempotent, butUPDATE products SET stock = <original_value>is. -
Reliability of Compensations: What happens if a compensating transaction fails? This is the "failure of a failure" scenario and is a hard problem.
- Solution 1: Retry. Use a persistent retry mechanism with exponential backoff and jitter.
- Solution 2: Alerting. If retries are exhausted, the failed compensation event must be moved to a dead-letter queue (DLQ) for manual intervention by an operations team. For a business process like an order, you cannot afford to leave the system in an inconsistent state (e.g., payment refunded but inventory not restocked).
-
Lack of Isolation: Sagas lack the isolation of ACID transactions. While one Saga is being compensated, another concurrent Saga might read the inconsistent intermediate state. For example, a user might see an item is out of stock, only for it to be restocked a moment later by a compensating transaction. This is a fundamental trade-off. You can mitigate this by using semantic locking (e.g., flagging an order as
CANCELLATION_IN_PROGRESS) or by accepting eventual consistency. -
Observability: As we've noted, tracing a choreographed flow is difficult. This is doubly true for rollbacks. Using a correlation ID that is passed through every event in the Saga (both forward and compensating) is non-negotiable. Structured logging with this ID allows you to reconstruct the entire lifecycle of a single business transaction across all services.
Conclusion
Today, we moved from the theory of Sagas to the practical implementation of their most critical component: compensation logic.
Key Takeaways:
- Compensation in a choreographed Saga is triggered by failure events published by downstream services.
- The implementation of a compensating transaction is the semantic inverse of the original local transaction and depends on its nature (DB update, API call, etc.).
- Idempotency is a mandatory property of compensating actions to handle retries safely.
- The failure of a compensating transaction is a serious risk that must be mitigated with robust retry mechanisms and, ultimately, a dead-letter queue for manual intervention.
- Achieving atomicity between state changes and event publishing, for both forward and compensating transactions, is typically accomplished using the Transactional Outbox pattern.
Preview of the Next Lesson:
We've spent the last few lessons deep in the weeds of a specific, powerful pattern for inter-service communication. Now, we'll zoom out to the strategic level of system design. Sagas are often needed when a business process crosses the boundaries between different parts of a domain. But how do we define those boundaries in the first place?
In our next lesson, we will begin a new module on Architectural Styles and Domain Modeling. Our first topic will be to design bounded contexts for a given business domain using DDD strategic patterns. This will provide the foundational architectural framework upon which patterns like Sagas are built.