The Cache Stampede Problem: Why a Popular Cache Key Can Take Down Your Backend
October 7, 2026
6 min read
When a frequently used cache entry expires under heavy traffic, every pending request can hit your database or origin service at the same instant. Here is why that happens and the techniques engineers use to stop it.
Ayodele Miracle JohnSOFTWARE ENGINEER
Caches are supposed to protect your backend from load. Most of the time they do. But there is a specific failure mode where the cache itself becomes the trigger for an outage: the cache stampede, also known as the thundering herd problem. It shows up in systems that otherwise look healthy, which is what makes it so disruptive when it finally happens.
What a Cache Stampede Is
A cache stampede happens when a single cached value, one that many requests depend on, expires or gets evicted, and a large number of requests notice the miss at roughly the same time. Instead of one request recomputing the value and everyone else reading it from cache, every one of those requests independently decides the cache is empty and goes to fetch or recompute the value itself.
If that value is expensive to produce (a complex database query, an aggregation across several services, a slow third-party API call), you now have dozens, hundreds, or thousands of copies of that expensive work running at once, all hitting the same downstream system within the same few hundred milliseconds. The system that the cache was shielding suddenly sees a spike that looks nothing like normal traffic, and it can fall over under the load.
The term "thundering herd" originally described a related operating systems problem, where many processes blocked on the same event all wake up at once when it fires, even though only one of them can actually make progress. The cache version is the same shape: one event (a cache miss) triggers many redundant reactions.
Why This Happens
The core issue is that a cache miss is usually interpreted the same way regardless of how many requests encounter it. Each request runs the same logic: check the cache, find nothing, recompute, write the result back. There is nothing in that logic that asks "is someone else already recomputing this right now?"
This is especially easy to trigger with a few common patterns:
A hot key with a hard TTL. If a widely used cache entry expires at a fixed time, and that key receives steady or bursty traffic, every request arriving in the moment right after expiry triggers its own cache miss.
A cold cache after a deploy or restart. Clearing or rebuilding a cache (intentionally or as a side effect of a deployment) empties out many hot keys at once. Traffic that was previously served entirely from cache now needs to be recomputed, all at the same time.
A popular key going viral. A sudden spike in requests for one specific item (a trending product page, a breaking news article, a suddenly popular API resource) can overwhelm the system the moment that item's cache entry needs to be refreshed, simply because the request volume for that one key is so much higher than normal.
A Concrete Example
Imagine an API endpoint that returns a leaderboard, computed from a query that scans a large table and takes 800 milliseconds to run. The result is cached for 60 seconds because the leaderboard does not need to be perfectly real time.
Under normal traffic, this works well: one request pays the 800 millisecond cost every 60 seconds, and everyone else reads the cached result in a few milliseconds. But suppose this endpoint gets hit 500 times per second. The instant the cached value expires, up to 500 requests in the next second can all see a cache miss before the first recomputation finishes and gets written back. That is potentially 500 concurrent 800 millisecond queries landing on the database at once, instead of one. Depending on the database's connection limits and query concurrency, this alone can be enough to cause timeouts, connection exhaustion, or a cascading slowdown that affects unrelated endpoints sharing the same database.
How Engineers Prevent Cache Stampedes
There is no single fix that applies everywhere, but a handful of techniques cover most real systems, and they are often combined.
Locking around recomputation. When a cache miss occurs, the first request to notice acquires a lock (often a short lived lock in Redis or a similar store) and does the work of recomputing the value. Other requests that see the same miss either wait briefly for the lock holder to finish and then read the fresh value, or fall back to a slightly stale value if one is available. This turns "many requests recompute" into "one request recomputes, everyone else waits or reads stale."
Request coalescing (single flight). Similar to locking, but implemented in application code: if a recomputation for a given key is already in flight, new requests for that same key are attached to the in-flight call instead of starting a new one, and all of them receive the same result when it completes. This avoids the lock and retry overhead of a distributed lock, at the cost of needing in-process coordination.
Jitter on TTLs. Instead of every instance of a cached value expiring at exactly the same time, add a small random offset to each TTL (for example, 60 seconds plus or minus 10 seconds, picked randomly per entry). This spreads expirations out over a window instead of letting them cluster, which reduces the odds that many keys, or many copies of the same key across cache nodes, expire in the same instant.
Stale while revalidate. Serve the existing cached value even after it has technically expired, while kicking off a background refresh to update it. Users get a fast response from slightly stale data, and only one request (or a small, bounded number) ever pays the cost of recomputing. This is the same idea used in HTTP's Cache-Control: stale-while-revalidate directive, and it generalizes well to application level caches too.
Probabilistic early expiration. Rather than waiting for a hard expiry, each read has a small, increasing probability of proactively triggering a background refresh as the entry approaches its expiry time. This spreads refresh work out over time and makes it very unlikely that an entry is ever read by many requests at the exact moment it goes stale. This approach was described in a well known paper from Facebook's engineering team and is sometimes referred to by the name used there, XFetch.
Choosing the Right Approach
Jitter is close to free and worth adding almost anywhere you have many cache entries with the same TTL, since it costs nothing but a small code change. Locking or request coalescing is the right call when recomputing the value is expensive and you can tolerate a short wait for the small number of requests that lose the race. Stale while revalidate is a good fit when slightly outdated data is acceptable and you want to avoid making any user wait on a slow recomputation. Probabilistic early expiration is worth the extra complexity for your highest traffic, highest cost cache entries, where even a brief stampede would be expensive.
None of these require tearing out an existing caching layer. Most caching libraries and services (Redis, Memcached, CDNs, and application level caches) support at least one of these patterns, and adding jitter or a simple lock around recomputation is often a small, isolated change rather than an architectural rewrite.
The broader lesson is that a cache's job is not just to store values, it is to make sure expensive work happens once, not once per concurrent request. Any caching strategy that does not account for what happens when many requests see the same miss at the same time is missing half the problem.

Oct 2, 2026
6 min read
The Circuit Breaker Pattern: Stopping Cascading Failures in Microservices
When one service in a distributed system starts failing, calls to it can pile up and take down everything upstream. The circuit breaker pattern stops that spread by failing fast instead of waiting.

Sep 30, 2026
5 min read
Rate Limiting Algorithms: Token Bucket vs Sliding Window Explained
A practical comparison of the two most common rate limiting algorithms, token bucket and sliding window, and guidance on which one fits a given API.

Sep 25, 2026
Why Your App Shows Stale Data Right After a Write: Understanding Read Replica Lag
Read replicas help databases scale, but they introduce a subtle bug where users don't see their own recent changes. Here's why that happens and how teams actually fix it.