Skip to main content
Create your own

Implementing Cache-Aside with Redis

Hello! Welcome back to our course on designing high-load distributed systems.

In the last lesson, we explored Redis eviction policies, which are crucial for managing the finite memory of your cache. We now know how to configure Redis to intelligently discard data when it's under memory pressure.

This lesson takes the next logical step: managing the flow of data into the cache from a primary data store. We will focus on one of the most fundamental and widely used caching strategies in system design.

The learning outcome for this lesson is to implement the cache-aside pattern for synchronizing Redis with a primary data store. This pattern, also known as "lazy loading," places the responsibility for managing the cache directly within your application logic. It's a versatile approach you've likely encountered, and today we'll formalize its implementation details, trade-offs, and failure modes.

1. The Cache-Aside Pattern: Core Logic

The core idea of cache-aside is that the cache sits "to the side" of the primary database. The application code acts as the orchestrator between the two, loading data into the cache on demand.

Let's start by examining the read path, which is where the "lazy loading" happens.

How to use Redis for Query Caching

This article from the official Redis documentation provides a clear, step-by-step explanation of the cache-aside pattern. It will walk you through the logic for both a cache miss and a cache hit.

Please read the sections titled 'Cache-aside with Redis (cache miss)' and 'Cache-aside with Redis (cache hit)'. Focus on the sequence of operations for each scenario.

As you've just read, the process is straightforward:

  1. The application receives a request for data.
  2. It first attempts to fetch this data from Redis.
  3. Cache Hit: If the data exists in Redis, it is returned directly to the client. The database is not involved.
  4. Cache Miss: If the data does not exist in Redis, the application queries the primary database (e.g., PostgreSQL, MongoDB).
  5. The application then stores the data retrieved from the database into Redis for subsequent requests.
  6. Finally, the data is returned to the client.

This flow is visualized in the diagram below.

This diagram illustrates the cache-aside read path. The 'Products Service' (your application) first queries Redis. On a miss, it queries the 'System of record' (the primary database), populates the cache, and then returns the data.

2. Implementing the Read Path

Now, let's move from theory to practice. A correct implementation requires careful consideration of a few key details: generating cache keys and managing data lifetime.

How to use Redis for Query Caching

This section of the Redis article provides a practical Node.js code example. It demonstrates how to generate a cache key and handle the cache miss logic.

Please read the section 'Implementing cache-aside with Redis and primary database'. Pay close attention to the getHashKey function and the logic within the router.post block, especially how the cache is populated on a miss.

From the example, we can extract two critical implementation details:

a. Cache Key Strategy

To avoid collisions and ensure you can retrieve the correct data, you need a deterministic way to generate a unique cache key from the request's parameters. The example uses a common strategy:

  • Serialize the query parameters: The req.body object, which contains the search filters, is serialized into a JSON string.
  • Hash the result: A SHA256 hash is computed from the string to create a consistent, fixed-length key. A prefix like CACHE_ASIDE_ is added to create a namespace for these keys within Redis.

This approach works well for complex query objects. For simpler key-value lookups (e.g., fetching a user by user_id), a simpler key like user:{user_id} is often sufficient and more debuggable.

b. Cache Expiration (TTL)

Notice this line in the code:
redis.set(hashKey, JSON.stringify(dbData), { EX: 60 });

Setting an expiration time (Time-To-Live or TTL) is crucial. It ensures that stale data doesn't remain in the cache forever. This is the simplest mechanism for cache invalidation. In the previous lesson, we discussed how policies like volatile-ttl or volatile-lru rely on TTLs to manage memory. Here, we see the application's role in setting those TTLs. The choice of 60 seconds is arbitrary; in a real system, this value would be tuned based on how long the data can afford to be stale.

3. Handling Writes: Synchronizing the Cache

The read path is only half the story. To keep the cache synchronized, we must handle data changes (creates, updates, and deletes). If we only implement the read path, the cache will become stale as soon as the underlying data is modified in the database.

The standard cache-aside approach for writes is cache invalidation.

Cache-aside (Lazy loading) Pattern

This article from System Design School provides excellent diagrams and a clear explanation of the write operation in the cache-aside pattern.

Please review the section 'How it works', focusing on the diagram and explanation for the 'Write operation'.

The write logic is:

  1. The application receives a write request (e.g., UPDATE, DELETE).
  2. The application writes the change directly to the primary database.
  3. Following a successful database write, the application issues a DEL command to Redis to remove the corresponding (now stale) cache entry.

Why not update the cache directly on a write?
This alternative, known as the write-through pattern, involves updating both the database and the cache. While it ensures the cache is always fresh, the cache-aside invalidation approach is often preferred because:

  • Simplicity: It's easier to implement DELETE than to re-calculate and SET the new value.
  • Efficiency: It avoids a "write-heavy, read-light" workload where data is written to the cache but never read again before being updated or evicted. The lazy-loading nature of cache-aside ensures only requested data occupies cache memory.

4. Advanced Considerations and Race Conditions

Given your experience with high-load systems, you know that the interaction between distributed components can lead to subtle race conditions. The cache-aside pattern is no exception.

Race Condition: Read vs. Write

Consider the following sequence of events:

  1. Request A (Read): Experiences a cache miss for key K.
  2. Request A (Read): Fetches the old value V1 from the database.
  3. Request B (Write): Updates the database value for key K to V2.
  4. Request B (Write): Invalidates key K in the cache by issuing a DEL.
  5. Request A (Read): Writes the stale value V1 it fetched in step 2 into the cache.

The result is that the cache now holds a stale value V1 that will persist until its TTL expires.

Mitigation Strategies:

  • TTL: This is the simplest and most common solution. A reasonably short TTL guarantees that the stale data will be purged eventually, limiting the window of inconsistency.
  • Write-through with locking: A more complex solution involves the write operation obtaining a short-lived lock on the cache key, updating the DB, and then updating the cache before releasing the lock. This adds significant complexity.
  • Lease Mechanism: On a cache miss, the read operation could place a "lease" (a temporary value with a very short TTL) in the cache. If a write operation sees this lease, it knows a read is in progress and can choose to wait or handle it differently. This is also complex and reserved for systems where strong consistency is paramount.

For most applications, accepting a small window of potential staleness and relying on a well-tuned TTL is the most practical trade-off.

Failure Mode: DB Write Succeeds, Cache Invalidation Fails

What happens if the database update succeeds, but the subsequent DEL command to Redis fails (e.g., due to a network partition)? The cache will now hold stale data indefinitely (or until TTL).

This is a classic distributed systems problem. A robust solution involves ensuring the cache invalidation happens reliably. The Transactional Outbox pattern, which we will cover in a future module, addresses this by atomically committing the database change and the intent to invalidate the cache in a single transaction. A separate process then reliably acts on that intent. For now, a common-but-imperfect approach is to implement retries with exponential backoff for the cache DEL operation.

Conclusion

In this lesson, we've implemented the cache-aside pattern, a cornerstone of caching architecture. You now understand not just the "happy path" but also the critical details of synchronization and the potential failure modes inherent in the pattern.

Key Takeaways:

  • Application-Managed: In the cache-aside pattern, the application code is responsible for orchestrating reads and writes to both the cache and the database.
  • Read Path (Lazy Loading): Read from cache -> On miss, read from DB -> Write to cache -> Return data.
  • Write Path (Invalidation): Write to DB -> Delete from cache.
  • Critical Details: A robust implementation depends on a deterministic cache key strategy and setting an appropriate TTL to manage staleness.
  • Trade-offs: The pattern's simplicity comes at the cost of potential race conditions and data inconsistency, which are typically managed with TTLs.

Preview of the Next Lesson:

Now that we have a solid pattern for interacting with Redis on a per-operation basis, our next lesson will focus on optimization. We will explore how to configure Redis pipelining and transactions for batching operations. This will allow us to reduce network overhead and perform multiple commands atomically, further improving the performance of our high-load system.

Can't find a good explanation? Sign up and we'll make it for you

Sign up