Your Redis Cache Is Fast… So Why Is Your Database Still Getting Destroyed?
Why a fast Redis cache can still overload PostgreSQL, and how to diagnose cache misses, stampedes, hot keys, TTLs, invalidation, and cache-aside failures.
8 min read
Redis Is Healthy. PostgreSQL Is Not.
The dashboard looks confusing:
Redis latency: 1.2 ms
Redis CPU: 18%
Cache hit rate: 91%
API requests: 10,000 / second
Database CPU: 94%
Database queries: 12,500 / secondRedis is fast. Most lookups are hits. Yet the database is still close to collapse.
The mistake is assuming that a fast cache is automatically an effective cache. Redis can answer every request quickly while the caching strategy sends a destructive amount of work to the database.
At 10,000 requests per second, a 9% miss rate is not small:
10,000 requests/sec × 9% misses = 900 database reads/secIf each miss triggers several queries, or thousands of requests miss the same key together, PostgreSQL receives the full blast.
Start With the Request Path
A typical cache-aside read looks like this:
Request
|
v
GET product:42 from Redis
|
+-- hit ---> return cached value
|
+-- miss ---> query PostgreSQL
|
v
SET Redis
|
v
return valueNothing is wrong with this pattern by itself. The failure appears when many requests take the miss path at once or when one miss performs much more database work than expected.
Failure 1: The Cache Stampede
Suppose product:42 is requested 4,000 times per second and its TTL expires.
Without coordination:
TTL expires
|
v
4,000 requests see a miss
|
v
4,000 requests query PostgreSQL
|
v
4,000 requests write the same Redis keyRedis remains fast because the misses are fast. PostgreSQL absorbs thousands of duplicate queries.
This is a cache stampede, sometimes called a thundering herd.
Use Request Coalescing
Only one request should rebuild a missing hot value. Other requests can briefly wait, use a stale value, or retry the cache.
async function getProduct(id: string) {
const key = `product:${id}`;
const cached = await redis.get(key);
if (cached) return JSON.parse(cached);
const lockKey = `lock:${key}`;
const ownsLock = await redis.set(lockKey, "1", { NX: true, PX: 3_000 });
if (!ownsLock) {
await new Promise((resolve) => setTimeout(resolve, 50));
const filled = await redis.get(key);
if (filled) return JSON.parse(filled);
throw new Error("Cache rebuild is still in progress");
}
try {
const product = await db.product.findUniqueOrThrow({ where: { id } });
await redis.set(key, JSON.stringify(product), { EX: 300 });
return product;
} finally {
await redis.del(lockKey);
}
}This simplified lock needs careful failure handling in a real system. The important property is that a miss does not create unlimited concurrent rebuilds.
Failure 2: Synchronized TTL Expiration
Imagine warming 100,000 keys during deployment with the same five-minute TTL:
12:00:00 100,000 keys written
12:05:00 100,000 keys expire
12:05:01 database traffic explodesAdd bounded random jitter so related keys do not expire simultaneously:
const baseTtlSeconds = 300;
const jitterSeconds = Math.floor(Math.random() * 60);
await redis.set(key, value, {
EX: baseTtlSeconds + jitterSeconds,
});Jitter spreads rebuild work across time. It does not fix bad queries or incorrect invalidation, but it prevents one timestamp from becoming a coordinated failure event.
Failure 3: A High Global Hit Rate Hides a Bad Endpoint
A 95% global hit rate can hide one critical route with a 20% hit rate.
Homepage cache hit rate: 99.8%
Product details hit rate: 98.0%
Personalized feed hit rate: 22.0%
Global hit rate: 95.1%If the personalized feed runs six expensive queries per miss, it may dominate database CPU even though the aggregate metric looks healthy.
Measure cache behavior by:
- Endpoint and operation
- Cache namespace
- Hit, miss, stale hit, and error
- Database queries triggered per miss
- Miss rebuild duration
- Key cardinality
- Tenant or traffic class where appropriate
The useful metric is not just cache_hit_total. It is the relationship between a miss and the downstream work that follows.
Failure 4: Hot Keys Move the Bottleneck
A hot key receives a disproportionate share of traffic.
product:42 30,000 reads/sec
product:108 120 reads/sec
product:901 14 reads/secThe key may overload a Redis shard, saturate network bandwidth, or create a severe stampede whenever it expires.
Possible responses include:
- Local in-process caching for very short-lived immutable data
- Replicating or sharding the value under multiple keys
- Stale-while-revalidate behavior
- Longer TTLs for data with tolerant freshness requirements
- Prewarming before a known traffic event
Every option changes consistency or operational complexity. Do not add local caching to frequently changing authorization data without understanding the stale-data risk.
Failure 5: Cache Invalidation Does Not Match Writes
Cache-aside reads are simple. Writes are where correctness becomes difficult.
Consider this order:
1. Update PostgreSQL
2. Delete Redis keyIf step 2 fails, stale data remains cached.
Now reverse it:
1. Delete Redis key
2. Update PostgreSQLA concurrent reader can miss after step 1, load the old database value, and put it back into Redis before step 2 commits.
There is no universally perfect invalidation sequence for every system. Common approaches include:
- Update the database, then invalidate the cache with retries
- Publish an invalidation event through an outbox
- Version cache keys so old values become unreachable
- Use short TTLs as a safety net
- Accept bounded staleness for read-heavy data
The right choice depends on whether the data is a product description, an account balance, a permission, or something else.
Diagnose Before Changing TTLs
During an incident, collect evidence in this order.
Check Cache Outcomes
cache_hits_total
cache_misses_total
cache_errors_total
cache_stale_hits_total
cache_rebuild_duration_msGraph them per endpoint, not only globally.
Check Redis Memory and Evictions
INFO memory
INFO statsLook for:
evicted_keysincreasing- Memory near
maxmemory - Unexpected key growth
- A policy that evicts important keys
If useful entries are constantly evicted, the application repeatedly falls through to the database even when TTLs appear reasonable.
Inspect Key Distribution
Use sampled application metrics or safe Redis tooling to identify hot namespaces and unexpectedly large keys. Avoid running blocking key scans such as KEYS * against a large production dataset.
Correlate Misses With Database Work
Compare the same timeline:
12:00 cache misses increase
12:00 DB queries/sec increase
12:01 DB CPU increases
12:02 API latency increasesThen inspect the database queries behind the misses. The existing guide on debugging a sudden database CPU spike covers that investigation in detail.
Choose the Right Caching Strategy
| Strategy | Strength | Main risk |
|---|---|---|
| Cache-aside | Simple and flexible | Miss stampedes and stale entries |
| Read-through | Centralizes loading behavior | Cache layer becomes more complex |
| Write-through | Cache updated with writes | Higher write latency and coupling |
| Write-behind | Fast writes | Data-loss and ordering risks |
| Stale-while-revalidate | Protects downstream systems | Readers may see bounded stale data |
Cache-aside is often a good default, but it still needs stampede protection, invalidation rules, observability, and a clear failure policy.
Production Best Practices
- Measure hit rate by endpoint and key namespace.
- Put strict bounds on concurrent cache rebuilds.
- Add TTL jitter for groups of related keys.
- Keep a database query efficient even when the cache misses.
- Treat Redis errors as a capacity event, not an invisible fallback.
- Decide whether stale data is safer than database overload.
- Monitor evictions, hot keys, memory, and rebuild latency.
- Load test cold-cache and mass-expiration scenarios.
- Use invalidation events or versioned keys where correctness demands it.
- Keep cache keys namespaced and ownership clear.
Conclusion
Redis latency tells you whether Redis is fast. It does not tell you whether your caching design protects the database.
The questions that matter are:
How many requests miss together?
What database work does one miss trigger?
Which keys are hot?
Why do entries disappear?
How is stale data invalidated?
What happens when Redis is unavailable?As traffic grows, these issues connect directly to what breaks between 100 and 10,000 requests per second. A production cache is not merely a fast key-value store. It is a controlled boundary between traffic and a more expensive dependency.
The caching architecture guide provides the broader design context for cache-aside, invalidation, stampede protection, hot keys, and Redis failure policies.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.