Skip to main content
Create your own
Lesson illustration

Circuit Breaker Pattern for Microservices

Hello! Welcome to our next lesson on building resilient systems.

In our last session, we focused on rate limiting, a crucial pattern for protecting your services from being overwhelmed by client requests. Today, we'll look at the other side of the coin: protecting your entire system from a single, failing service. This is where the Circuit Breaker pattern comes in. It's a fundamental concept for anyone building microservice architectures, and a common topic in senior-level system design interviews.

Your goal for this lesson is to learn how to implement the Circuit Breaker pattern to prevent these cascading failures. We'll explore the theory behind it, its different states of operation, and then dive into a practical, production-grade implementation in Go.

The Problem: Cascading Failures

In a distributed system, services often depend on one another. For example, a UserService might call an AuthService, which in turn might query a Database. What happens if the Database becomes slow or unresponsive? The AuthService's requests start to time out, its connection pools fill up, and it eventually becomes unresponsive too. This failure then "cascades" up to the UserService, which also becomes unresponsive. Soon, a single downstream failure can bring down your entire application.

This chain reaction is called a cascading failure. The video below provides a great, concise explanation of this problem and introduces why a circuit breaker is needed.

Introduction to circuit breaker in microservices (for beginners)

This video from sudoCODE explains what a circuit breaker is in the context of microservices and the problems it solves, like cascading failures.

Watch the first part of the video, from the beginning to failure is returned. Focus on how a failure in one service can spread to others.

The image below illustrates this concept visually. On the left, you see how a failure in Service C propagates upstream, causing timeouts and failures in Services B and A. On the right, you see the solution: the Circuit Breaker intercepts the calls, preventing the system from repeatedly hitting a known-failing service.

This diagram contrasts a cascading failure scenario with the protective routing provided by a Circuit Breaker.

The Solution: A Three-State Machine

A Circuit Breaker acts as a proxy for operations that might fail, such as network calls to another service. It operates as a state machine with three states: Closed, Open, and Half-Open. The goal is to "fail fast" by blocking calls to a failing service, giving it time to recover.

This diagram clearly shows the three states of the Circuit Breaker pattern and the transitions between them.

Let's break down each state, using the diagram above as our guide:

  1. Closed State:

    • This is the normal, healthy state.
    • Requests pass through to the downstream service as usual.
    • The circuit breaker monitors the outcomes of these requests, counting successes and failures.
    • If the number of failures (based on a rate, consecutive count, etc.) exceeds a configured threshold, the circuit breaker "trips" and transitions to the Open state.
  2. Open State:

    • The circuit breaker has detected a fault.
    • All subsequent requests are immediately rejected ("short-circuited") without attempting to contact the downstream service. This is the "fail fast" principle in action.
    • This prevents the calling service from being blocked and gives the failing service "breathing room" to recover.
    • A cooldown timer starts. When this timer expires, the circuit transitions to the Half-Open state.
  3. Half-Open State:

    • This is a cautious, probing state.
    • The circuit breaker allows a limited number of "trial" requests to pass through to the downstream service.
    • If these trial requests succeed, the circuit breaker determines that the service has recovered. It transitions back to the Closed state and resets its failure counts.
    • If any trial request fails, the circuit breaker assumes the service is still unhealthy. It immediately transitions back to the Open state, and the cooldown timer starts again.

Implementing a Circuit Breaker in Go

Now let's move from theory to practice. Since you're proficient in Go, we'll use sony/gobreaker, a popular and mature library, to implement a circuit breaker. This library makes it straightforward to wrap your service calls and configure the breaker's behavior.

This article from OneUptime provides an excellent, comprehensive walkthrough. We'll use it to guide our implementation.

How to Implement Circuit Breakers in Go with sony/gobreaker

This article introduces the sony/gobreaker library and shows a basic implementation.

First, read the introduction, which covers the core concepts and states we just discussed. Then, review the basic implementation to see how to install the library and wrap a function call with cb.Execute().

Fine-Tuning with Configuration

The real power of a circuit breaker lies in its configuration. These settings are not arbitrary; they are critical design decisions that you must tune based on the specific service's characteristics and your system's Service Level Objectives (SLOs).

The gobreaker.Settings struct lets you control this behavior.

How to Implement Circuit Breakers in Go with sony/gobreaker

This section of the article details the key configuration options for tuning the circuit breaker's behavior.

Read the section on Comprehensive Configuration Options. Pay close attention to these key settings: Timeout: The cooldown period for the Open state. How long should you wait before trying again? ReadyToTrip: A function that decides when to trip from Closed to Open. This is where you define your failure threshold (e.g., based on failure ratio or consecutive failures). MaxRequests: The number of trial requests to allow in the Half-Open state. OnStateChange: A callback function that's triggered on any state change. This is invaluable for logging and monitoring.

Distinguishing Failure Types

A crucial aspect of a robust circuit breaker is distinguishing between different types of errors. For example, an HTTP 404 Not Found error is not a service failure; it's a valid response indicating a client error. Tripping the circuit for such errors would be incorrect. However, an HTTP 503 Service Unavailable or a connection timeout is a service failure and should contribute to tripping the circuit.

The IsSuccessful setting in gobreaker allows you to define this logic. The article provides an excellent example.

How to Implement Circuit Breakers in Go with sony/gobreaker

This section shows how to implement custom logic to decide which errors should count as failures.

Read the section on Custom Failure Detection. The isRecoverableError function demonstrates exactly how to ignore client-side HTTP status codes (4xx) while treating server-side issues (5xx) and other errors as genuine failures.

This same principle applies directly to gRPC, which you've worked with before. The article "How to Implement Circuit Breakers for gRPC Services" shows an identical IsSuccessful function, but for gRPC status codes.

How to Implement Circuit Breakers for gRPC Services

This excerpt shows how the same error-filtering logic applies to gRPC.

In this article, find the gobreaker example and look at the IsSuccessful function. Notice how it checks for gRPC codes like codes.InvalidArgument and codes.NotFound and considers them successes (i.e., not service failures), just as we did for HTTP 4xx codes.

Beyond Failing: Fallback Strategies

Simply failing fast when a circuit is open is better than cascading failure, but we can do better. A resilient system should degrade gracefully. This means having a fallback strategy for when the primary operation is unavailable.

What you do as a fallback depends heavily on the business context of the service.

How to Implement Circuit Breakers in Go with sony/gobreaker

This section outlines three essential fallback strategies with Go code examples.

Read the section on Implementing Fallback Strategies. Focus on understanding these three distinct approaches: Static Fallback Response: If you can't get live data, can you return a default value or, even better, a slightly stale value from a cache? Graceful Degradation: If a non-critical feature (like a recommendation engine) is unavailable, can you still serve the main content of the page without it? This often involves using the OnStateChange callback to toggle a feature flag. Queue for Retry: For background tasks (like sending a confirmation email), if the operation fails, can you place it in a queue to be retried later when the circuit closes?

Having well-defined fallback strategies is a hallmark of a senior engineer's approach to system design.

Conclusion

Today, we've explored the Circuit Breaker pattern, a powerful tool for building resilient, fault-tolerant distributed systems. You've learned how it prevents cascading failures and how to implement and configure it in Go for practical, real-world scenarios.

Key Takeaways:

  • Purpose: Circuit Breakers prevent cascading failures by isolating failing services and allowing them to recover.
  • Three States: The pattern operates in a Closed (healthy), Open (faulted, fail-fast), or Half-Open (probing) state.
  • Configuration is Crucial: The thresholds for tripping, cooldown timeouts, and trial request counts must be tuned to your service's specific needs.
  • Smart Failure Detection: A robust implementation distinguishes between client errors (e.g., HTTP 4xx) and true service failures (e.g., HTTP 5xx, timeouts).
  • Always Have a Fallback: A circuit breaker should be paired with a fallback strategy—like serving cached data, degrading gracefully, or queuing for retry—to improve the user experience.

In our next lesson, we will begin a new module on Infrastructure and Observability. We'll start by looking at how we distribute traffic across our services in the first place by diving into the world of load balancing and exploring the critical differences between Layer 4 and Layer 7 load balancers.

Can't find a good explanation? Sign up and we'll make it for you

Sign up