Skip to main content
Create your own

Client-Side Circuit Breaker Implementation

Hello! Welcome back to our module on Resilience and Failure Handling Patterns.

In our previous lesson, we explored the Retry pattern with exponential backoff and jitter. This pattern is your first line of defense, designed to handle short-lived, transient failures by transparently re-issuing a request. However, this approach assumes the downstream service will recover quickly.

Today, we address a different scenario: what happens when a service is completely unavailable or severely degraded for a longer period? Continuously retrying in this situation is wasteful. It consumes client-side resources, adds latency to user requests, and can even hinder the struggling service's recovery.

This brings us to our learning outcome for this lesson: to implement the Circuit Breaker pattern using a client-side library. This pattern acts as a stateful "fuse," allowing a system to detect a sustained failure, stop sending requests, and gracefully degrade functionality.

1. The Circuit Breaker: A State Machine for Resilience

The Circuit Breaker pattern, popularized by Michael Nygard in his book Release It!, is a state machine that wraps protected function calls and monitors them for failures. It prevents an application from repeatedly attempting an operation that is likely to fail.

Much like an electrical circuit breaker, it has three primary states:

  • CLOSED: The default state. Requests are allowed to pass through to the downstream service. The circuit breaker monitors for failures. If the failure rate exceeds a configured threshold, the breaker "trips" and moves to the OPEN state.
  • OPEN: The breaker has tripped. For a configured duration (the waitDurationInOpenState), all requests to the downstream service are immediately rejected with an exception, without any attempt to execute them. This is the "fail-fast" behavior. After the wait duration elapses, the breaker transitions to HALF_OPEN.
  • HALF_OPEN: The breaker allows a limited number of "trial" requests to pass through. If these requests succeed, the breaker concludes the service has recovered and transitions back to CLOSED. If they fail, it assumes the problem persists and transitions back to OPEN, starting the wait timer again.

To get a precise understanding of this state machine and its behavior, we'll use the documentation for Resilience4j, a popular and lightweight fault tolerance library for Java. Your background in applied mathematics and physics will make the finite state machine model intuitive.

CircuitBreaker

This documentation for Resilience4j provides a clear, technical overview of the Circuit Breaker's implementation. It formally defines the states and the logic for transitioning between them.

Please read the 'Introduction' to understand the state machine concept. Then, read the sections starting with 'Failure rate and slow call rate thresholds' and 'The CircuitBreaker rejects calls...'. Focus on how failure thresholds trigger the transition to OPEN, and how the HALF_OPEN state facilitates recovery.

2. Monitoring Failures: The Sliding Window

A key question is how the circuit breaker determines that the failure rate has crossed a threshold. It doesn't react to a single failure; instead, it aggregates the outcomes of recent calls using a sliding window. This approach provides a more statistically sound basis for tripping the circuit.

Resilience4j supports two types of sliding windows:

  1. Count-based Window: Aggregates the outcome of the last N calls. For example, if N=100, it always considers the success/failure status of the most recent 100 calls.
  2. Time-based Window: Aggregates the outcomes of calls within the last N seconds. For example, if N=60, it considers all calls that occurred in the last minute.

The choice between them depends on the nature of your traffic. A count-based window is suitable for consistent traffic, while a time-based window can be more appropriate for bursty traffic patterns.

The following reading details how these windows are implemented. Given your experience with data structures and algorithms, you'll appreciate the discussion of their performance characteristics.

CircuitBreaker

Let's continue with the Resilience4j documentation to understand how call outcomes are measured.

Please read the sections 'Count-based sliding window' and 'Time-based sliding window'. Note the implementation details (circular arrays) and the time/space complexity mentioned for each.

3. Implementing the Pattern with Resilience4j

Now we'll move from theory to a practical implementation. Using a library like Resilience4j is highly recommended over building your own, as it provides a thread-safe, configurable, and battle-tested solution.

The process involves three main steps:

  1. Configure the circuit breaker's parameters.
  2. Decorate the function call you want to protect.
  3. Provide a fallback mechanism for when the circuit is open.

Step 1: Configuration

This is where you define the behavior of your circuit breaker. The configuration is critical and must be tuned based on the specific service being called and your availability requirements.

The Resilience4j documentation provides a comprehensive guide to its configuration options.

CircuitBreaker

This section of the documentation is the most important for practical implementation. It details all the knobs you can turn to customize the circuit breaker's behavior.

Read the section 'Create and configure a CircuitBreaker'. Pay close attention to the configuration table and the builder example. Focus on understanding these key parameters: failureRateThreshold, waitDurationInOpenState, slidingWindowType, slidingWindowSize, minimumNumberOfCalls, and permittedNumberOfCallsInHalfOpenState.

As you can see, you have fine-grained control. For example, in a high-load payment system, you might configure recordExceptions to treat only network-related or timeout exceptions as failures, while ignoring BusinessException types (like "Invalid CVV") that don't indicate a service health issue.

Step 2: Decorating a Function Call

Once configured, you apply the circuit breaker to your code. Resilience4j uses a functional approach, allowing you to "decorate" any Supplier, Callable, or other functional interface.

CircuitBreaker

This part of the documentation shows how to wrap your business logic with the circuit breaker you've configured.

Please read the section 'Decorate and execute a functional interface'. The code snippet shows how a CheckedFunction0 (a supplier that can throw a checked exception) is decorated. This is the core mechanism for applying the pattern.

Step 3: Integration and Fallbacks in a Real Application

Decorating a single function is straightforward. In a real application, especially one using a framework like Spring, the integration is often more seamless. Let's look at an example from the Spring Cloud Circuit Breaker guide, which uses Resilience4j as its underlying implementation.

This guide demonstrates wrapping a call made with WebClient (a modern, reactive HTTP client) and providing a fallback. The fallback is the logic that executes when the circuit is open, allowing the application to degrade gracefully—for example, by returning a cached response or a default value.

Getting Started | Spring Cloud Circuit Breaker Guide

This guide from spring.io demonstrates how to integrate a circuit breaker into a Spring Boot application. It provides a practical view of how the pattern is used in a microservices context.

Please review the code snippets in the sections 'Apply The Circuit Breaker Pattern' (starting from 'Spring Cloud Circuit Breaker provides an interface...') and the updated ReadingApplication.java. Focus on how ReactiveCircuitBreakerFactory is used to create a ReactiveCircuitBreaker and how the .run() method is used to wrap the WebClient call and provide a fallback Mono.

In the BookService example, the run method takes two arguments:

  1. The primary action: webClient.get().uri("/recommended")...
  2. The fallback function: throwable -> Mono.just("Cloud Native Java (O'Reilly)")

This is a powerful composition. If the WebClient call fails enough to open the circuit, subsequent calls to bookService.readingList() will immediately execute the fallback function, returning a default book recommendation without attempting the network call.

4. Observability: Knowing Your Circuit's State

A silent circuit breaker is not very useful. For monitoring and debugging, it's essential to know when a circuit changes state. Resilience4j provides a rich eventing system that allows you to log state transitions, successful calls, errors, and more.

As an engineering manager, you know that metrics and observability are non-negotiable for production systems. This is how you would hook into the circuit breaker's lifecycle to feed data into your monitoring stack (e.g., Prometheus, Grafana).

CircuitBreaker

Finally, let's see how to monitor the circuit breaker's activity.

Briefly review the section 'Consume emitted CircuitBreakerEvents'. Note the different event types (onSuccess, onError, onStateTransition) you can subscribe to. This is the key to observability.

Conclusion

In this lesson, we have moved beyond simple retries to a more robust, stateful resilience pattern. The Circuit Breaker is a critical tool for building systems that can withstand partial failures and prevent them from cascading.

Key Takeaways:

  • State Machine: The Circuit Breaker operates as a state machine with CLOSED, OPEN, and HALF_OPEN states to control request flow.
  • Fail-Fast: In the OPEN state, it immediately rejects requests, protecting the downstream service and preventing the client from wasting resources.
  • Sliding Windows: It uses count-based or time-based sliding windows to make informed decisions about when to trip, based on recent failure rates.
  • Library Implementation: Using a library like Resilience4j provides a configurable, thread-safe, and observable implementation, which can be cleanly integrated into application code using decoration and fallbacks.

Preview of the Next Lesson

We've now implemented the Circuit Breaker pattern within our application's code using a client-side library. However, this is not the only way. Modern infrastructure, particularly service meshes, offers another approach where this logic is moved out of the application and into the infrastructure layer.

In the next lesson, we will configure circuit breaking in a service mesh proxy (e.g., Envoy) and analyze the trade-offs versus a library-based approach. This will provide a broader perspective on implementing resilience patterns in a distributed architecture.

Can't find a good explanation? Sign up and we'll make it for you

Sign up