Skip to main content
Create your own

Implementing Retry with Exponential Backoff and Jitter

Hello! Welcome to the first lesson in our module on Resilience and Failure Handling Patterns.

Over the past modules, we've explored the core components of distributed systems, from network infrastructure and load balancers to various data stores like PostgreSQL, MongoDB, and Kafka. Now, we'll shift our focus to the patterns that ensure these components can communicate reliably and gracefully handle the inevitable failures that occur in a distributed environment.

Today, we'll start with a fundamental building block of system resilience. Our goal is to implement the Retry pattern with exponential backoff and jitter. This pattern is a first line of defense against the transient faults that are common in any large-scale system.

1. The Case for Retries and Their Inherent Risks

In any distributed system, requests can fail for temporary reasons: a brief network partition, a server process restarting, or a momentary resource overload. Instead of immediately propagating the failure to the user, the client can simply try the request again. This is the essence of the Retry pattern.

However, naive retries can be dangerous. If a service is already struggling with high load, a wave of immediate retries from all its clients can push it over the edge, creating a feedback loop that leads to a full-blown outage.

To understand these dynamics, let's start with an article from the AWS Builders' Library, which provides excellent context based on Amazon's extensive operational experience.

Timeouts, retries and backoff with jitter

This article, 'Timeouts, retries and backoff with jitter,' clearly explains the rationale behind these resilience patterns. It sets the stage by discussing why retries are necessary and introduces the significant risks they pose if not implemented carefully.

Please read the introductory section 'Failures Happen' and the following section 'Retries and backoff'. Focus on the concept of retries being 'selfish' and the example of how load can be amplified in a multi-layered architecture.

As you read, consider your experience with high-load payment systems. A key takeaway from the article is the importance of idempotency. An operation is idempotent if it can be performed multiple times without changing the result beyond the initial application. In a financial system, retrying a non-idempotent "charge credit card" API could lead to multiple charges, making idempotency a critical design constraint.

2. A Smarter Delay: Exponential Backoff

To prevent clients from overwhelming a struggling service, we need a more intelligent delay strategy than simply waiting a fixed amount of time. The most common approach is exponential backoff. The core idea is to increase the wait time exponentially after each failed attempt. This gives a struggling service progressively more time to recover.

The following article provides a clear, code-driven explanation of this strategy.

Mastering Exponential Backoff in Distributed Systems

The article 'Mastering Exponential Backoff in Distributed Systems' from Betterstack offers a practical guide to implementation. We'll use it to break down the mechanics of exponential backoff.

Please read the sections '1. Retry mechanism' and '2. Backoff strategy: Timing is everything'. Pay attention to the formula for exponential backoff and the JavaScript implementation provided.

The standard formula for the delay is:

For example, with a base_delay of 100ms and a factor of 2, the delays would be 100ms, 200ms, 400ms, 800ms, and so on.

In practice, this exponential growth is usually capped at a maximum value (max_delay) to prevent unreasonably long waits, especially for user-facing requests.

3. Breaking Synchronization: The Role of Jitter

Exponential backoff is effective, but it has a subtle flaw. If multiple clients experience a failure at the same time, they will all back off and retry in synchronized waves. This can cause periodic spikes of traffic, a phenomenon known as the thundering herd problem.

The solution is to introduce randomness, or jitter, to the backoff delay. This spreads out the retry attempts over time, smoothing the load on the downstream service.

Let's continue with the same resources to understand jitter.

Timeouts, retries and backoff with jitter

We'll now examine how adding jitter solves the thundering herd problem. The AWS article provides the 'why,' and the Betterstack article provides the 'how.'

First, read the 'Jitter' section in the AWS article to solidify your understanding of the problem.

Mastering Exponential Backoff in Distributed Systems

Now, let's look at a concrete implementation of jitter.

Read the section '3. Jitter: Avoiding the thundering herd'. Note the different jitter strategies mentioned, particularly 'Full Jitter,' which is often the most effective.

With full jitter, the delay is a random value between 0 and the calculated exponential backoff delay.

actual_delay = random(0, exponential_backoff_delay)

This strategy effectively breaks the synchronization between clients. The following graph visually demonstrates the impact of jitter on system load.

This graph compares the total number of calls (work) made to a service by competing clients using different backoff strategies. Notice how the 'FullJitter' and 'EqualJitter' strategies result in significantly less work as client contention increases, compared to plain 'Exponential' backoff or no backoff ('None'). This illustrates how jitter reduces wasteful retries and improves overall system efficiency.

4. A Practical, Reusable Implementation

Now, let's combine these concepts—retry limits, exponential backoff, jitter, and max delays—into a robust, reusable implementation. A production-ready implementation must also be able to distinguish between errors that are safe to retry and those that are not.

For instance, a 503 Service Unavailable error or a network timeout (ETIMEDOUT) are good candidates for a retry. However, a 400 Bad Request or 401 Unauthorized error will almost certainly fail again with the same request, so retrying is wasteful.

The Betterstack article provides an excellent, encapsulated implementation in a JavaScript class. Let's analyze its structure.

Mastering Exponential Backoff in Distributed Systems

This final reading provides a complete, configurable class for handling retries. It's a great template for a production-ready implementation.

Please review the section 'Implementing exponential backoff'. Focus on the structure of the ExponentialBackoff class, its configurable parameters, and especially the isRetryable method. This method contains the critical logic for deciding when a retry is appropriate.

This class-based approach is powerful because it encapsulates the complex retry logic and provides a clean interface (.execute(operation)) to the application code. The configurability of maxRetries, initialDelay, maxDelay, and the list of retryableErrors allows you to tune the behavior for different use cases within your system.

Conclusion

In this lesson, we've dissected one of the most critical patterns for building resilient distributed systems. You are now equipped to implement it effectively.

Here are the key takeaways:

  • Retry Pattern: A fundamental technique for handling transient failures by re-issuing a failed request.
  • Exponential Backoff: A strategy to prevent retries from overwhelming a service by exponentially increasing the delay between attempts.
  • Jitter: The addition of randomness to backoff delays to prevent synchronized retries (the "thundering herd" problem) and smooth out load.
  • Idempotency and Retryable Errors: A robust implementation requires careful consideration of which operations are safe to retry (idempotent operations) and which errors indicate a transient, retryable fault.

Preview of the Next Lesson

Retries are excellent for short-lived, transient issues. But what if a downstream service is completely unavailable for an extended period? In that scenario, continuing to send requests—even with backoff—is wasteful and can hinder the service's ability to recover.

In our next lesson, we will address this by exploring the Circuit Breaker pattern. This pattern allows a client to detect a prolonged failure, stop sending requests for a period, and gracefully handle the unavailability of a dependency.

Can't find a good explanation? Sign up and we'll make it for you

Sign up