Skip to main content
Create your own

API Gateway Rate Limiting with Token Bucket

Introduction

In our last lesson, we dove deep into the core of the inference engine, designing a sophisticated scheduling policy to manage stateful, multi-turn tool-calling. You learned how Continuum's TTL-based KV cache retention makes an economic trade-off between latency and throughput, a critical optimization for agentic workloads.

Today, we move from optimizing the performance of individual requests inside the engine to protecting the stability and fairness of the service as a whole. Your learning outcome for this lesson is to implement token-bucket rate limiting at the API gateway level.

This is a fundamental pillar of any production-grade service. Without it, a single misbehaving client or a sudden burst of traffic could overwhelm your carefully tuned inference engine, degrading performance for all users and potentially incurring huge costs. We'll explore why rate limiting is especially critical for LLM APIs and then implement a robust, industry-standard solution.

Why Standard Rate Limiting Fails for LLMs

First, let's understand why the unique nature of LLM inference demands a more nuanced approach than traditional rate limiting.

For a typical API, where each request represents a similar amount of work (e.g., a database lookup), a simple Requests Per Minute (RPM) limit is often sufficient. However, for an LLM API, the computational cost of requests is highly variable:

  • A request with a short prompt and a 10-token response is cheap.
  • A request with a 4,000-token context and a 2,000-token generation is extremely expensive in terms of both memory (for the KV cache) and compute (for token decoding).

Limiting by request count is therefore unfair and inefficient. A user could exhaust their quota with a few heavy requests, starving many users with small, quick queries.

To learn more about the specific challenges and best practices for LLM rate limiting, please read the following sections from an article by Hivenet.

Rate limiting and quotas for LLM APIs

This article, 'Rate limiting and quotas for LLM APIs', directly addresses the shortcomings of traditional request-based limits and proposes token-aware alternatives.

Please read the sections titled 'Why plain request limits fail for LLMs', 'Token-aware patterns that work', and 'Choosing your unit: requests vs tokens'. Focus on the arguments for using Tokens Per Minute (TPM) as a more equitable and cost-aligned metric.

As the article highlights, the most effective approach is token-aware rate limiting, typically measured in Tokens Per Minute (TPM). This aligns the limit directly with the resource consumption, providing a much fairer and more predictable way to manage load.

Choosing an Algorithm: The Case for Token Bucket

Now that we've established what to measure (tokens), we need to decide how to measure and enforce the limit. There are several algorithms, but the token bucket algorithm strikes an excellent balance of simplicity, effectiveness, and memory efficiency.

Token Bucket Algorithm Explained
This diagram illustrates the core concept of the token bucket algorithm. A bucket has a certain capacity and is refilled with tokens at a constant rate. Each incoming request (or in our case, each token to be processed) consumes one or more tokens from the bucket. If the bucket is empty, the request is rejected.

The key advantages of the token bucket algorithm are:

  • Handles Bursts: It allows for short bursts of traffic up to the bucket's capacity, which is ideal for the naturally "bursty" nature of user interactions.
  • Enforces Sustained Rate: Over the long term, the throughput is limited by the constant refill rate.
  • Memory Efficient: For each user or API key, you only need to store two values: the current number of tokens and the timestamp of the last refill.

To understand how token bucket solves the problems of simpler algorithms like the fixed-window counter, please watch this short video.

Rate Limiter System Design: Token Bucket, Leaky Bucket, Scaling

The video 'Rate Limiter System Design' from ByteByteGo provides a clear, concise explanation of the token bucket algorithm and why it's superior to the naive fixed-window approach.

Watch from the beginning to 04:30. Pay attention to how the fixed window counter can lead to double the intended traffic at window boundaries, and how the token bucket's refill mechanism elegantly solves this.

Architectural Placement: The API Gateway

With the algorithm chosen, the next critical decision is where to place the rate-limiting logic. The most common and effective location is at the API gateway or as a middleware layer in your web framework.

LLM Serving System Architecture with Gateway and Rate-Limiting
This architectural diagram shows an 'LLM Gateway' as the entry point for requests. Placing the rate limiter here ensures that it can inspect and reject traffic *before* it consumes resources from the core model-serving components.

Placing the rate limiter at the gateway has several advantages:

  1. Separation of Concerns: It keeps the cross-cutting concern of rate limiting out of your core application logic.
  2. Protection: The inference service is shielded from traffic storms, as rejected requests never reach it.
  3. Centralized Control: A single point of enforcement simplifies policy management.

Implementing a Distributed Rate Limiter with Redis

Now for the practical part. Since our service might run on multiple server instances, we can't store the token buckets in the local memory of each gateway. We need a centralized, shared, and fast data store. Redis is the industry standard for this task.

For each rate-limited entity (e.g., an API key), we'll store its token bucket state in a Redis hash. However, this introduces a classic distributed systems problem: a race condition.

Consider this sequence:

  1. Gateway A receives a request for key xyz. It reads the token count from Redis: 10.
  2. Simultaneously, Gateway B receives another request for key xyz. It also reads the count: 10.
  3. Gateway A decides the request is allowed, decrements the count to 9, and writes it back to Redis.
  4. Gateway B, unaware of Gateway A's action, also decides the request is allowed, decrements its stale value to 9, and writes it back to Redis.

The result: two requests were processed, but the token count was only decremented once. We've allowed excess traffic.

The solution is to perform the entire "read-calculate-update" sequence as a single atomic operation. Redis provides a powerful mechanism for this: Lua scripting. A Lua script sent to Redis is guaranteed to execute atomically, preventing any other command from running concurrently.

To understand this implementation pattern in detail, please read the following sections from the "Design a Distributed Rate Limiter" article.

Design a Distributed Rate Limiter

This article from Hello Interview provides a superb, practical guide to implementing a distributed rate limiter, focusing on the token bucket algorithm with Redis.

Please read the following two sections carefully: Start at the heading 'For our system, we'll go with the Token Bucket algorithm.' and read until the heading '3) When limits are exceeded...'. This section explains how to use Redis for state and, crucially, describes the race condition and its solution using a Lua script. Next, read the section titled '3) When limits are exceeded, reject requests with HTTP 429 and helpful headers'. This covers the standard practice for how your API should respond when a limit is hit.

Implementation Sketch in Python

Let's translate this into a practical implementation sketch for a FastAPI middleware. This code demonstrates how to use a Lua script with a Redis client to perform atomic rate limiting.

Here is the Lua script that encapsulates the token bucket logic. It calculates the number of tokens to add since the last request, refills the bucket, and then checks if a token can be consumed.

-- rate_limiter.lua

local key = KEYS[1]            -- The Redis key for the user, e.g., 'rate_limit:user123'
local bucket_capacity = tonumber(ARGV[1]) -- Max tokens the bucket can hold
local refill_rate = tonumber(ARGV[2])     -- Tokens to add per second
local cost = tonumber(ARGV[3])            -- Tokens this request costs (e.g., 1 for RPM, or N for TPM)

local current_time = redis.call('TIME')
local now = tonumber(current_time[1]) + tonumber(current_time[2]) / 1000000

local bucket_state = redis.call('HMGET', key, 'tokens', 'last_refill_ts')
local current_tokens = tonumber(bucket_state[1])
local last_refill_ts = tonumber(bucket_state[2])

if current_tokens == nil then
  current_tokens = bucket_capacity
  last_refill_ts = now
else
  local time_elapsed = now - last_refill_ts
  local tokens_to_add = time_elapsed * refill_rate
  current_tokens = math.min(bucket_capacity, current_tokens + tokens_to_add)
  last_refill_ts = now
end

local allowed = 0
if current_tokens >= cost then
  current_tokens = current_tokens - cost
  allowed = 1
end

redis.call('HMSET', key, 'tokens', current_tokens, 'last_refill_ts', last_refill_ts)
redis.call('EXPIRE', key, 3600) -- Expire the key after 1 hour of inactivity

return {allowed, bucket_capacity, current_tokens, math.ceil(cost / refill_rate)}

And here is the Python middleware that would load and execute this script for each incoming request.

import redis.asyncio as redis
import time
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse




# --- Configuration ---
REDIS_HOST = "localhost"
REDIS_PORT = 6379
BUCKET_CAPACITY = 100  # Max burst
REFILL_RATE = 10       # Tokens per second

app = FastAPI()
redis_client = None




# Load the Lua script from file
with open("rate_limiter.lua", "r") as f:
    LUA_SCRIPT = f.read()
    
lua_sha = None

@app.on_event("startup")
async def startup_event():
    global redis_client, lua_sha
    redis_client = await redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)



    # Load script into Redis and get its SHA hash for efficient future calls
    lua_sha = await redis_client.script_load(LUA_SCRIPT)

@app.middleware("http")
async def rate_limiting_middleware(request: Request, call_next):



    # For simplicity, we use the client's host as the identifier.
    # In a real app, you'd use an API key or user ID from the Authorization header.
    client_id = request.client.host
    key = f"rate_limit:{client_id}"




    # For TPM, you'd estimate the token cost here. We'll use a fixed cost of 1 for RPM.
    request_cost = 1


```grasp
{
  "type": "exercise",
  "id": "9d7d1b04-def5-4eb9-b74f-a952a73d815f"
}
# Execute the Lua script atomically
results = await redis_client.evalsha(
    lua_sha, 1, key, BUCKET_CAPACITY, REFILL_RATE, request_cost
)

allowed, limit, remaining, retry_after = [int(r) for r in results]

headers = {
    "X-RateLimit-Limit": str(limit),
    "X-RateLimit-Remaining": str(remaining),
    "Retry-After": str(retry_after)
}

if not allowed:
    return JSONResponse(
        status_code=429,
        content={"error": "Rate limit exceeded"},
        headers=headers
    )

response = await call_next(request)



# Add headers to successful responses too
response.headers.update(headers)
return response

@app.get("/")
async def root():
return {"message": "Hello World"}


This implementation provides a robust, production-ready foundation for rate limiting. To adapt it for **TPM**, the `request_cost` would need to be dynamically calculated based on the prompt's token count, which could be done in the middleware before calling the Lua script.




### Conclusion

In this lesson, you've designed and implemented a critical component for any production LLM service. By moving from the engine's internals to the API gateway, you've learned how to protect your entire system from overuse and ensure fair resource allocation.

**Key Takeaways:**
*   **LLMs require token-aware rate limiting (TPM)** because the cost per request is highly variable.
*   The **Token Bucket algorithm** is an industry standard that gracefully handles bursts while enforcing a sustained rate.
*   Rate limiting should be implemented at the **API gateway level** to separate concerns and protect backend services.
*   Using a centralized store like **Redis** is necessary for distributed systems, and **Lua scripting** is the key to performing atomic read-modify-write operations to prevent race conditions.

**Preview of the Next Lesson:**
With this lesson, you have added another crucial piece to your mental model of a complete LLM serving stack. We are now approaching the culmination of the first half of this course. In the next module, you will begin your **Capstone Project: Build a Custom Inference Engine.** You will start by designing the high-level architecture, bringing together the concepts we've covered—from the API server and scheduler to the execution engine and cache manager—into a coherent system design.


Can't find a good explanation? Sign up and we'll make it for you

Sign up