Skip to main content
Create your own

Nginx Retry Policies: Balancing Availability and Load

Hello! Welcome back.

In our last lesson, we established how to use timeouts in nginx to detect failures and protect your system from unresponsive upstream services. Detecting a failure is the first step; reacting to it intelligently is the next.

This lesson focuses on building a more resilient system by automatically handling transient failures. We will explore how to configure nginx to retry failed requests, a powerful technique for improving availability. However, as with any powerful tool, it comes with significant risks if misused.

By the end of this 60-minute lesson, you will be able to configure retry policies and budgets in nginx for upstream services and analyze the trade-offs between availability and load. We'll cover the specific directives for implementing retries, discuss the danger of "retry storms," and examine advanced patterns like retry budgets for mitigating this risk.

1. Configuring Basic Retry Policies in Nginx

When an upstream server fails—perhaps due to a temporary network glitch, a brief overload, or a pod restarting—simply failing the client's request might be premature. If you have multiple backend servers, nginx can transparently retry the request on a different server, potentially masking the transient error from the end-user.

Let's look at the core nginx directives that control this behavior.

Avoiding the Top 10 NGINX Configuration Mistakes

To begin, let's look at a practical example of how retry logic is configured within an upstream block. This article from the NGINX blog provides a concise and clear demonstration.

Please read the section 'Mistake 10: Not Taking Advantage of Upstream Groups'. Focus on the proxy_next_upstream directive in the example configuration. Note which conditions are defined to trigger a retry and how it works within the upstream group context.

As the article demonstrates, the key to enabling retries is the proxy_next_upstream directive. It allows you to define a set of conditions that nginx considers a "failure," prompting it to pass the request to the next server in the upstream group.

For a more exhaustive list of these conditions and the directives that limit retries, another resource is helpful.

GEP-1731: HTTPRoute Retries

The Kubernetes Gateway API Enhancement Proposal on retries includes a background section that neatly summarizes the relevant NGINX directives.

Please read the 'NGINX' subsection under 'Background on implementations'. Pay close attention to the list of possible values for proxy_next_upstream and the purpose of proxy_next_upstream_tries and proxy_next_upstream_timeout.

Based on these resources, we can summarize the main components of an nginx retry policy:

  • proxy_next_upstream: Defines the failure conditions. Common values include:

    • error: A network error occurred when connecting to or communicating with the server.
    • timeout: A timeout was reached (as defined by proxy_connect_timeout, proxy_read_timeout, etc.).
    • http_500, http_502, http_503, http_504: The upstream returned one of these specific HTTP error codes. These are often signs of a temporary server-side issue.
    • non_idempotent: Allows retrying non-idempotent methods like POST or PATCH. This is extremely dangerous and should be used with caution. Retrying a POST request to a payment endpoint, for example, could result in a double charge. It's generally safe only if your backend has its own idempotency handling mechanism (e.g., using an Idempotency-Key header).
  • proxy_next_upstream_tries: Limits the maximum number of attempts for a request. A value of 3 means one initial attempt plus two retries. A value of 0 disables the limit.

  • proxy_next_upstream_timeout: Sets a time limit during which retries can be performed. This prevents a request from being retried indefinitely if the upstream servers are flapping. A value of 0 disables this time limit.

Here is a sample configuration incorporating these directives:

upstream backend_service {
    zone upstreams 64K;
    # List of servers in the pool
    server 10.0.1.10:8080;
    server 10.0.1.11:8080;
    server 10.0.1.12:8080;
}

server {
    listen 80;
    location / {
        proxy_pass http://backend_service;
        proxy_connect_timeout 2s;

        # Retry policy
        proxy_next_upstream error timeout http_503;
        proxy_next_upstream_tries 3;
        proxy_next_upstream_timeout 10s;
    }
}

In this setup, if a request to one server fails with a network error, a timeout, or an HTTP 503, nginx will try the next server, up to a total of 3 attempts, as long as the total time spent does not exceed 10 seconds.

A Critical Distinction: Retries vs. Health Checks

It's important to differentiate the proxy_next_upstream mechanism from the max_fails and fail_timeout parameters on the server directive.

  • proxy_next_upstream is about the current request. It decides whether to give this specific request another chance on a different server.
  • max_fails / fail_timeout is about the health of the server. It's a passive health check mechanism. If a server fails max_fails times within fail_timeout seconds, nginx marks it as "down" and stops sending any new requests to it for the fail_timeout duration.

These two mechanisms work together. proxy_next_upstream handles immediate, per-request recovery, while max_fails provides a longer-term circuit breaker to isolate a consistently failing node.

2. The Trade-Off: Availability vs. Load Amplification

Retries are a double-edged sword. While they improve resilience against transient faults, they can be catastrophic during a broader service degradation.

The Upside: Increased Availability
Imagine a service where 0.1% of requests fail due to random network hiccups. Without retries, your error rate is 0.1%. With a single retry, the probability of two consecutive independent failures is 0.1% * 0.1% = 0.0001%. You've just improved your availability by three orders of magnitude for this class of error.

The Downside: The "Retry Storm"
Now, consider a different scenario. A backend service is overloaded due to a spike in traffic or a buggy deployment. It starts responding slowly, causing proxy_read_timeout errors, or it returns 503 Service Unavailable.

What does our retry policy do? For every incoming request that fails, it generates two more.

A retry storm amplifies load on an already failing system. A single incoming request can trigger multiple retry attempts, turning a slowdown into a complete outage.

This creates a dangerous positive feedback loop:

  1. Service becomes slow/overloaded.
  2. Nginx sees timeouts or 5xx errors.
  3. Nginx retries the requests, multiplying the load on the service.
  4. The increased load makes the service even slower, causing more failures.
  5. The cycle repeats, leading to a "retry storm" or "thundering herd" that can cause a complete outage.

This amplification effect is a classic pattern of cascading failure in distributed systems. Your experience with high-load payment and trading systems has likely shown how quickly such feedback loops can destabilize a complex environment.

3. Advanced Mitigation: Retry Budgets

To prevent retry storms, we need to make our retry logic smarter. Instead of retrying blindly, the system should stop retrying when the overall error rate becomes too high. This is the principle behind a retry budget.

A retry budget limits the volume of retries to a fraction of the volume of successful requests. For example, you might set a budget that "the number of retried requests per second cannot exceed 20% of the number of successful requests per second."

  • During normal operation: The error rate is low. Failed requests are retried, and the budget is not exceeded. Availability is improved.
  • During a major incident: The error rate spikes. The retry budget is quickly exhausted. The system stops retrying and instead "fails fast," returning errors directly to the client. This sheds load from the struggling backend, preventing a complete collapse and giving it a chance to recover.

This mechanism introduces a stabilizing negative feedback loop.

It is crucial to note that Nginx Open Source does not have a built-in directive for retry budgets. This is an advanced feature found in other proxies like Envoy or service meshes like Istio. However, understanding the concept is vital for designing robust high-load systems.

The GEP-1731 resource you reviewed earlier provides a good conceptual overview of budgets, as it proposes adding them to the Kubernetes Gateway API, drawing inspiration from Envoy's implementation.

In Envoy, a retry budget is typically configured with two parameters:

  • budget_percent: The ratio of retries to total requests allowed. A common default is 20%.
  • min_retry_concurrency: A floor for the number of retries allowed, ensuring that even with very low traffic, a few retries can still happen.

While you can't configure this with a simple nginx directive, you could theoretically implement similar logic using Lua scripting within nginx, but that adds significant complexity. For systems where this level of control is needed, migrating the retry logic to a more advanced proxy like Envoy is often the more practical path.

Conclusion

In this lesson, we've moved from simply detecting failures with timeouts to actively recovering from them with retries.

Key Takeaways:

  • Retry Configuration: Nginx uses proxy_next_upstream to define what constitutes a retriable failure, proxy_next_upstream_tries to limit the number of attempts, and proxy_next_upstream_timeout to limit the duration.
  • Idempotency is Key: Only retry requests that are idempotent (safe to be repeated). Retrying non-idempotent requests (POST, PATCH) without a backend idempotency mechanism can lead to data corruption.
  • The Core Trade-Off: Retries increase availability during transient, isolated failures but risk causing "retry storms" that amplify load and create cascading failures during widespread service degradation.
  • Retry Budgets: Advanced proxies use retry budgets to prevent storms by limiting the ratio of retries to successful requests, creating a stabilizing feedback loop that prioritizes system stability over individual request success during an outage.
  • Tooling: Basic retry policies are easily configured in nginx. Advanced patterns like retry budgets require more sophisticated tools like Envoy or custom Lua scripting.

In our next lesson, we will broaden our perspective from handling failures within a single upstream group to managing traffic at a larger scale. We will configure nginx as an edge proxy to distribute traffic across simulated datacenter clusters using weighted routing, taking our first step into multi-datacenter architectures.

Can't find a good explanation? Sign up and we'll make it for you

Sign up