Caching Architecture: Cache-Aside, Write-Through, Write-Behind, and Redis Patterns
Learn how to design caching with Redis using cache-aside, write-through, write-behind, invalidation, TTLs, stampede protection, hot-key handling, and production failure strategies.
34 min read
The Cache Made the API Faster — Then Returned the Wrong Data
Imagine a product endpoint takes:
180 msbecause every request reads from PostgreSQL.
We add Redis.
Now the same endpoint takes:
8 msGreat.
Then a customer updates a product price.
PostgreSQL contains:
price = 120but Redis still contains:
price = 100The API reads Redis first and returns:
100So now we have:
Database: 120
Redis: 100
User sees: 100The API became faster.
But the data became stale.
That is the main trade-off of caching.
A cache is another copy of your data.
And whenever you create another copy, you must answer:
How does this copy get created?
How long may it live?
How does it become stale?
How is it updated or removed?
What happens if Redis fails?
What happens when thousands
of requests miss at once?Caching is not just:
Redis before databaseIt is a consistency and load-management decision.
Why Do We Cache?
The usual reason is performance.
Without cache:
Request
↓
Database
↓
ResponseSuppose database latency is:
150 msWith Redis:
Request
↓
Redis
↓
Responsemaybe:
5 msCaching can also reduce:
database CPU
database connections
expensive queries
external API calls
network trafficBut now we have two copies:
Authoritative data
→ database
Temporary copy
→ cacheThe database is usually the source of truth.
Redis is usually a faster derived copy.
A Simple Mental Model
Think of caching as:
Source of truth
↓
temporary faster copyFor example:
PostgreSQL
↓
RedisThe cache exists because reading the original source is more expensive.
The difficult question is:
How do we keep that faster copy useful without letting it become dangerously stale?
Different caching patterns answer that question differently.
Main Caching Patterns
The common strategies are:
Cache-Aside
Read-Through
Write-Through
Write-BehindThey mainly differ in:
who loads the cache
when writes reach the cache
when writes reach the databaseCache-Aside
Cache-aside is probably the most common Redis pattern.
The application manages the cache directly.
The read flow is:
Request
↓
Redis
├── HIT
│ ↓
│ return
│
└── MISS
↓
database
↓
cache result
↓
returnIn simpler words:
Check cache.
If found:
return it.
If not found:
load from database,
store in Redis,
then return.Cache-Aside Example
async function getProduct(id: string) {
const key = `product:${id}:v1`;
const cached = await redis.get(key);
if (cached) {
return JSON.parse(cached);
}
const product = await db.product.findUnique({
where: { id },
});
if (!product) {
return null;
}
await redis.set(key, JSON.stringify(product), {
EX: 300,
});
return product;
}The flow is:
GET product
↓
Redis lookup
↓
hit?
┌──┴──┐
yes no
↓ ↓
return PostgreSQL
↓
cache result
↓
returnWhy Cache-Aside Is Popular
It is simple.
The application decides:
what to cache
which key to use
how long it should live
when to invalidate itIt also avoids caching data nobody reads.
Suppose there are:
10 million productsbut only:
100,000are frequently requested.
Cache-aside naturally keeps only hot data in Redis.
The Downside: The First Request Is Still Slow
Suppose the key does not exist.
The request still needs:
Redis miss
+
database query
+
Redis writeSo the first request might take:
180 mswhile later requests take:
8 msThis is called a:
cold cacheor:
cache missCache Hit vs Cache Miss
A cache hit means:
requested data exists in cacheFlow:
Request
→ Redis
→ value found
→ returnA cache miss means:
requested data does not existFlow:
Request
→ Redis
→ missing
→ database
→ Redis
→ returnOne of the most important cache metrics is therefore:
hit rateHit Rate Example
Suppose:
10,000 requestsand:
9,500 hits
500 missesHit rate:
9,500 / 10,000
= 95%That sounds excellent.
But a global hit rate can hide dangerous endpoints.
We will return to that later.
The Hard Part: What Happens After a Write?
Suppose Redis currently contains:
product:42
price = 100Then the customer updates the product:
price = 120Database becomes:
120But Redis may still contain:
100If we do nothing, users continue seeing stale data.
This is where cache invalidation starts.
A Common Write Strategy: Database First, Then Delete Cache
Example:
const product = await db.product.update({
where: { id },
data: input,
});
await redis.del(`product:${id}:v1`);
return product;The flow:
Update database
↓
Commit succeeds
↓
Delete Redis keyThe next read becomes:
Redis miss
↓
load fresh DB value
↓
repopulate RedisWhy Delete Instead of Updating the Cache?
You might think:
Database updated to 120
↓
Set Redis to 120That can work.
But deletion is often simpler.
When we delete:
product:42the next read rebuilds the value from the authoritative database.
Conceptually:
write path
→ invalidate
read path
→ rebuildThis reduces the chance that write logic and cache-building logic produce different representations.
Why Update the Database First?
Consider the opposite:
Update Redis first
↓
Update databaseWhat if:
Redis succeeds
Database failsNow:
Redis: 120
Database: 100The API may return a value that never actually committed.
Database-first is usually safer when the database is the source of truth.
But Database-Then-Delete Is Still Not Perfect
Consider:
Database update succeeds
↓
process crashes
↓
Redis DEL never happensNow:
Database: 120
Redis: 100The stale cache remains.
This is why TTLs are important even when you actively invalidate.
TTL as a Safety Net
Suppose the key is created with:
EX: 300;That means:
expire after 300 secondsor:
5 minutesSo even if invalidation fails:
stale value survives
for at most about 5 minutesThis makes TTL a safety mechanism.
It is not only a memory-management setting.
What Is TTL?
TTL means:
Time To LiveExample:
product:42
TTL = 300 secondsAfter 300 seconds Redis automatically removes the key.
Then the next request reloads it from the database.
Choosing TTL
There is no universal TTL.
Different data can tolerate different staleness.
For example:
Public product description
→ 10 minutes may be fine
Product price
→ maybe 30 seconds
Feature configuration
→ maybe 1 minute
User permissions
→ perhaps very short,
or do not cache stale values at all
Stock inventory
→ maybe cache only carefullyThe first question should be:
How stale may this data safely become?
Then choose the caching policy.
Strong Correctness vs Faster Reads
Suppose:
marketing pageshows slightly old product text.
Maybe acceptable.
But suppose:
authorization decisionuses a stale cache and grants access that should have been revoked.
That is much more serious.
Not all data deserves the same caching strategy.
Cache Invalidation Is a Distributed-Systems Problem
PostgreSQL and Redis are separate systems.
We cannot normally do:
BEGIN TRANSACTION
update PostgreSQL
delete Redis
COMMIT BOTHas one simple atomic transaction.
That means there is always a failure window.
For example:
DB succeeds
Redis failsor:
DB commits
process crashes
before cache invalidationCaching architecture must decide how much inconsistency is acceptable and how recovery works.
Reliable Invalidation With an Outbox
Suppose updating a product must reliably invalidate a cache projection.
Instead of:
update DB
↓
best-effort Redis DELwe can write an invalidation event in the same database transaction.
For example:
Database transaction
├── update product
└── insert outbox event
↓
commitThen a worker processes:
ProductUpdatedand runs:
DEL product:42:v1If Redis is temporarily unavailable, the event can retry.
This is more reliable than an untracked background promise.
But Even Invalidation Has Races
Consider this sequence.
Initial database:
price = 100Cache is empty.
Then:
Reader
→ cache miss
Reader
→ queries DB
→ gets 100Before Reader stores it:
Writer
→ updates DB to 120
Writer
→ deletes cache keyThen Reader continues:
Reader
→ writes old value 100
into cacheNow:
Database = 120
Redis = 100The writer correctly deleted the cache.
But the concurrent reader recreated stale data afterward.
This is the classic delete race.
Delete Race Timeline
Reader: cache miss
Reader: SELECT
gets price 100
Writer: UPDATE
price = 120
Writer: DEL cache
Reader: SET cache
price = 100Final state:
DB = 120
Cache = 100This demonstrates an important fact:
There is no universal Redis invalidation sequence that magically makes Redis transactional with PostgreSQL.
How Can We Reduce the Delete Race?
Possible strategies include:
short TTL
versioned cached values
rebuild coordination
reliable invalidation events
routing critical reads to DBWhich one to use depends on correctness requirements.
Version Cache Entries
Suppose database records have:
version = 42A cache entry can include:
{
"version": 42,
"price": 120
}If an older rebuild tries to write:
version = 41the cache layer can reject it or avoid replacing a newer version.
This adds complexity but helps with concurrent rebuild races.
Versioned Cache Keys
Another common approach:
product:42:v1Then later the representation changes:
product:42:v2Now old serialized data does not conflict with the new format.
This kind of versioning is useful for schema evolution.
What Versioned Keys Do Not Automatically Solve
Changing:
v1
→ v2helps with:
cache schema changes
serialization changes
large-scale invalidationIt does not automatically solve:
concurrent stale reads
write races
data correctnessThose are different concerns.
Write-Through Caching
With write-through, the application writes through a cache abstraction.
Conceptually:
Application
↓
Cache layer
├── database
└── cacheThe write path updates both.
For example:
update product
↓
cache abstraction
↓
database update
↓
cache updateAfter the write, the cache is already warm.
Why Use Write-Through?
Suppose updated products are almost always read immediately.
With database-then-delete:
write
↓
delete cache
↓
next request misses
↓
database againWith write-through:
write
↓
database + cache
↓
next read hitsFrequently read data remains warm.
Write-Through Trade-Offs
Writes now pay extra latency.
Without cache update:
database writeWith write-through:
database write
+
Redis writeMore importantly, we still have two systems.
Suppose:
PostgreSQL succeeds
Redis failsWhat should happen?
Possible choices:
return success
and recover cache lateror:
fail requestor:
retry cache updateThere is no automatic atomicity just because we call the pattern "write-through."
Write-Through Does Not Mean Distributed Transaction
This is important.
Redis
+
PostgreSQLare still independent systems.
Write-through describes:
when cache update happensnot:
perfect atomic consistencyYou still need an explicit failure policy.
Write-Behind Caching
Write-behind changes the model significantly.
Instead of synchronously persisting every change to the database:
Request
↓
cache / buffer
↓
success response
↓
database write laterThe application may acknowledge before the database has the final value.
Why Write-Behind Can Be Fast
Suppose the database supports:
10,000 writes/secbut incoming traffic temporarily reaches:
50,000 writes/secA write-behind buffer may absorb bursts.
The client writes quickly to:
Redis / durable log / queueand workers persist to the database later.
This can smooth write traffic.
Write-Behind Changes the Source of Truth Temporarily
Immediately after success:
Redis/buffer
→ contains latest value
Database
→ may still contain old valueThat means the application must understand delayed persistence.
This is a much stronger architectural choice than ordinary caching.
Never Implement Write-Behind as a Background Promise
Bad:
await redis.set(...);
void db.product.update(...);
return success;What if the process crashes after returning success?
The database write disappears.
The user was told:
successbut durable state never changed.
That is data loss.
Durable Write-Behind
If you use write-behind, use something durable.
For example:
Request
↓
durable queue/log
↓
success
↓
worker
↓
databaseNow if a worker crashes, the write still exists in the queue.
But this introduces new concerns:
ordering
retries
duplicates
backpressure
dead-letter handling
replay
lagWrite-behind is not just a cache option.
It becomes part of the data-write architecture.
When Is Write-Behind Appropriate?
Possible fits:
analytics counters
telemetry
high-volume aggregates
non-critical activity
batched writesLess appropriate for:
payments
account balances
critical inventory decisions
security changesunless the system is specifically designed around durable asynchronous state.
Read-Through Caching
Cache-aside puts miss logic in application code:
app
→ cache
→ DB on missRead-through moves that behavior into a cache abstraction or provider.
The application asks:
get product 42The abstraction handles:
cache hit
or
load DB on missConceptually:
Application
↓
Data abstraction
├── Redis
└── DBWhy Use Read-Through?
It can centralize:
TTL policy
serialization
miss loading
cache metricsand remove repeated cache-aside code from services.
Read-Through Trade-Off
Abstraction can hide cost.
The application sees:
getProduct(id);but sometimes it is:
5 ms Redis hitand sometimes:
300 ms expensive DB querySo even when caching is hidden behind a library, observe:
hit rate
miss rate
load duration
downstream query costCache Stampede
One of the most dangerous caching failures happens when a hot key expires.
Suppose:
product:42receives:
5,000 requests/secNormally:
all requests
→ Redis hitThen its TTL expires.
Suddenly:
5,000 requests
→ cache miss
→ databaseInstead of one DB query, you may get:
5,000 DB queriesfor the same value.
This is called a:
cache stampedeor:
thundering herdStampede Timeline
Before expiry:
Request × 5000
↓
Redis
↓
one cached valueAfter expiry:
Request × 5000
↓
Redis miss
↓
PostgreSQL × 5000A cache that normally protects the database suddenly directs enormous traffic toward it.
Request Coalescing
One solution is:
one request rebuilds
everyone else waitsFor example:
5,000 misses
↓
one rebuild lock
↓
1 DB query
↓
populate cache
↓
serve waiting requestsThis is often called:
single flight
request coalescing
mutex rebuildThe Lock Must Also Be Bounded
Do not make 100,000 requests wait forever behind one rebuild.
If the rebuild is slow or failing, waiting requests may still exhaust:
memory
sockets
request deadlinesUse bounded waiting and deadlines.
If necessary:
serve stale value
or
rejectinstead.
Stale-While-Revalidate
Another useful strategy is:
serve slightly stale value
while one worker refreshes itExample:
Cache value:
fresh for 60 sec
Then:
stale-but-usable for 30 secDuring stale window:
request
→ gets old value immediately
background refresh
→ fetches latest DB valueThis avoids sending every request to the database when freshness expires.
When Stale-While-Revalidate Is Good
Useful for:
public content
catalog data
configuration
search metadata
recommendationswhere:
slightly staleis better than:
slow or unavailableDo not use stale data blindly for:
authorization
payment decisions
critical balancesTTL Jitter
Suppose you cache:
100,000 product keysall with exactly:
TTL = 300 secondsThey were created during one deployment.
Five minutes later:
100,000 keys expire togetherNow the database receives a massive spike.
Instead use TTL jitter.
For example:
base TTL = 300 sec
actual TTL =
270–330 secNow expirations spread over time.
TTL Jitter Example
function ttlWithJitter(baseSeconds: number) {
const jitter = Math.floor(Math.random() * 60) - 30;
return baseSeconds + jitter;
}Then:
300becomes values like:
274
312
297
329This reduces synchronized expiration.
Proactive Refresh
For a few known extremely hot keys, waiting for expiration may be unnecessary.
Suppose:
homepage:featured-productsreceives:
100,000 reads/secInstead of letting it expire, refresh before expiry.
For example:
TTL = 10 minutes
refresh job
→ every 8 minutesThis keeps the hot key warm.
Do Not Depend on Proactive Refresh Alone
Refresh jobs can fail.
So still keep:
TTL
miss behavior
fallbackdefined.
Otherwise one failed scheduler can break the cache strategy.
Negative Caching
Suppose clients repeatedly request:
/product/does-not-existWithout negative caching:
request
→ Redis miss
→ DB
→ not found
request
→ Redis miss
→ DB
→ not found
request
→ Redis miss
→ DB
→ not foundThis can generate large database traffic.
Cache "Not Found"
Example:
const NOT_FOUND = "__not_found__";
if (!product) {
await redis.set(key, NOT_FOUND, {
EX: 30,
});
return null;
}Next request:
Redis
→ NOT_FOUND
→ return 404/nullwithout another DB query.
Keep Negative TTL Short
Suppose:
product 123does not exist.
We cache:
not foundfor:
1 hourThen product 123 is created one minute later.
Clients may continue seeing:
not foundfor another 59 minutes.
So negative caching usually uses a shorter TTL.
For example:
15–60 secondsdepending on the domain.
Negative Caching Can Be Abused
An attacker sends:
/random-id-1
/random-id-2
/random-id-3
...If every miss creates a Redis key:
millions of negative-cache entriescan consume memory.
Control:
key cardinality
TTL
identity validation
request rateNegative caching should not allow attackers to turn Redis into unlimited storage.
Hot Keys
Suppose one key:
homepage:configreceives:
50,000 requests/secEven Redis is fast, but every request still requires:
network
Redis command processing
serialization
deserializationOne Redis shard may become overloaded.
This is a:
hot keyLocal In-Process Cache
For small slowly changing data, you can add:
Request
↓
local memory cache
↓
Redis
↓
databaseSuppose each Node.js instance caches:
homepage configfor:
5 secondsThen thousands of local requests may require only one Redis read every few seconds per instance.
This reduces Redis load.
Local Cache Trade-Off
Now there are three copies:
Database
Redis
Node process memoryEach Node instance may hold a different value temporarily.
Example:
Instance A
→ config v10
Instance B
→ config v11until local TTL expires.
So local caches should generally use:
short TTLsand only for data where this staleness is acceptable.
Avoid Local Cache for Security-Sensitive State
Do not casually locally cache:
revoked permissions
account lock status
security policy
critical access decisionsbecause every process has its own stale copy.
For correctness-sensitive decisions, use a source with the required consistency guarantee.
Cache Key Design
Bad keys:
42
user7
configThey provide little context.
Better:
catalog:product:42:v3or:
tenant:acme:permissions:user-7:v2A good cache key should communicate:
owner
resource
identity
schema/versionKey Namespace Example
catalog:
product:
42:
v3Actual key:
catalog:product:42:v3This makes operational debugging easier.
Why Include a Version?
Suppose cached JSON used to be:
{
"id": 42,
"price": 100
}Then the application changes to:
{
"id": 42,
"priceMinor": 10000,
"currency": "USD"
}Old cached JSON may break the new application.
Change key:
product:42:v1to:
product:42:v2Now old entries are ignored naturally.
Avoid Wildcard Deletion as the Main Strategy
Suppose you create:
product:42:*and later scan Redis to delete everything matching it.
Large wildcard operations can be expensive and difficult to reason about.
Prefer known deterministic keys where possible.
For grouped invalidation, consider:
version namespaces
generation counters
explicit key registriesdepending on the architecture.
Value Size Matters
Redis is fast, but large values cost:
memory
network bandwidth
serialization CPU
deserialization CPUCaching:
50 KB objectand caching:
20 MB objectare very different choices.
Measure actual value sizes.
Avoid Huge Cached Lists
Example:
all 5 million productsin one Redis JSON value.
Problems:
large network transfer
expensive serialization
all-or-nothing invalidation
hot key
memory usageCache smaller projections where possible.
Compression
Compression can reduce:
Redis memory
network bytesbut increases:
CPU
latency
implementation complexityIt may help for larger values.
Measure before enabling it everywhere.
Redis Failure Is a Production Design Question
Suppose normal traffic is:
10,000 requests/secCache hit rate:
95%Database normally receives:
500 requests/secThen Redis fails.
Now potentially:
10,000 requests/secfall through to PostgreSQL.
Database capacity:
1,000 requests/secResult:
Redis fails
↓
database overloads
↓
whole service failsThe cache outage becomes a database outage.
"Just Fall Back to the Database" Can Be Dangerous
This sounds resilient:
if Redis fails:
read from DBBut at scale it may create an avalanche.
The database was sized assuming the cache absorbed most traffic.
So the miss path must also be capacity-limited.
Bound Database Fallback Concurrency
Example:
Redis unavailable
↓
fallback gate
↓
maximum 100 DB queries
↓
databaseOther requests may:
serve stale data
wait briefly
or
receive 503depending on the endpoint.
This protects the database.
Short Redis Timeouts
Suppose Redis is degraded and every call waits:
5 secondsNow requests pile up.
Use a short Redis deadline.
For example:
10–100 msdepending on network architecture and workload.
If Redis cannot answer quickly, trigger the defined fallback.
Different Endpoints Can Have Different Redis Failure Policies
For public catalog data:
Redis failure
→ maybe use short-lived local stale copyFor permissions:
Redis failure
→ maybe query authoritative DB
with strict bounded concurrencyFor optional recommendations:
Redis failure
→ skip recommendationsThere does not need to be one global policy.
Caching Is Also Load Shaping
Caching does more than reduce latency.
It controls where requests go.
Normally:
10,000 requests
↓
9,500 Redis
500 databaseIf Redis misses suddenly:
10,000 requests
↓
databaseSo the cache architecture controls traffic distribution.
Cache Hit Rate Is Not Enough
Suppose global cache hit rate is:
95%Looks excellent.
But imagine:
GET /products/:id
→ 99% hit rate
→ cheap miss
GET /reports/dashboard
→ 50% hit rate
→ miss runs a 5-second queryThe second endpoint may still overload PostgreSQL.
So measure by:
endpoint
cache namespace
operationnot just globally.
Measure the Cost of a Miss
A miss might cause:
one indexed SELECTor:
15 joins
4 external calls
large aggregationThese misses are not equal.
A useful metric is:
database work caused by cache missor at least:
miss latency
rebuild latencyCache Freshness Is Also a Metric
Suppose users are receiving cached values.
It can be useful to know:
How old is this data?For example:
cachedAt = 12:00
servedAt = 12:07
data age = 7 minutesA cache can have a fantastic hit rate while serving unacceptable staleness.
Important Cache Metrics
Measure:
hit rate
miss rate
stale-hit rate
Redis latency
Redis errors
Redis timeouts
rebuild duration
rebuild lock wait
database fallback rate
evictions
key count
memory usage
value size
network throughput
hot keys
data ageRedis Eviction Matters
Suppose Redis memory is full.
Depending on configuration, Redis may evict keys.
Then:
hot key unexpectedly disappears
↓
cache miss
↓
database rebuildIf many keys are evicted:
miss rate rises
database load risesMonitor:
evicted_keysand Redis memory pressure.
Expiration vs Eviction
These are different.
Expiration:
TTL reached
→ key intentionally removedEviction:
Redis memory pressure
→ key removed to free spaceA sudden rise in eviction can look like a cache-stampede problem even though TTLs did not expire.
Cache Cardinality
Suppose each user creates:
20 cache keysand you have:
5 million usersPotential state:
100 million keysEven tiny values can become expensive.
Track:
number of keys by namespacenot only total Redis memory.
Choose What Should Be Cached
Good candidates often have some of these properties:
read frequently
change less frequently
expensive to compute
safe to be slightly stale
reused across requestsExamples:
product metadata
public configuration
expensive aggregate
reference data
feature metadataPoor Cache Candidates
Caching may provide little value when:
data changes constantly
every request is unique
cache hit probability is low
correctness requires latest value
values are huge
database query is already extremely cheapAdding Redis to every endpoint can make the architecture slower and more complicated.
Cache User Profile?
Suppose:
GET /meis requested frequently.
Could cache it.
But ask:
What fields change?
How often?
Are permissions included?
How quickly must changes appear?
Can different devices update it?Maybe cache:
display name
avatar URL
preferencesbut not:
security-sensitive authorization stateunder the same long TTL.
Cache design can be field-specific.
Cache Settings?
Suppose:
system settingsare read on every request and updated rarely.
That may be a strong cache candidate.
Example:
Request
→ local 5-sec cache
→ Redis
→ databaseWhen admin updates settings:
DB commit
→ invalidate RedisLocal copies expire quickly.
This can dramatically reduce database reads.
Cache Sessions?
Redis is commonly used for session storage, but this is slightly different from ordinary caching.
If Redis holds the only active session state, it may be:
primary operational statenot merely a disposable cache.
That changes:
durability
replication
failure
backup
availabilityrequirements.
Do not assume every Redis use case is a cache.
Cache vs Source of Truth
A useful question is:
If Redis loses this key permanently, can the system reconstruct it?
If yes:
likely cache/derived stateIf no:
Redis may be holding authoritative stateand should be designed differently.
A Practical Cache-Aside Flow
For a product read:
GET /products/42
↓
build key:
catalog:product:42:v1
↓
Redis GET
↓
hit?
┌────┴────┐
yes no
↓ ↓
return acquire rebuild lock
↓
check cache again
↓
database
↓
Redis SET
TTL + jitter
↓
release lock
↓
returnThe second cache check after acquiring the lock is important because another request may already have rebuilt the value.
Why Double-Check After the Lock?
Example:
Request A misses
Request B missesA gets rebuild lock.
A loads DB and populates Redis.
B then gets the lock.
Without another cache check:
B loads DB againWith another check:
B sees A's cached value
→ returns immediatelyThis prevents unnecessary rebuilds.
A Practical Write Flow
Suppose a product changes.
A practical design may be:
PATCH /products/42
↓
validate
↓
database transaction
↓
commit product
↓
publish ProductUpdated
through outbox
↓
returnThen:
outbox worker
↓
DEL catalog:product:42:v1Meanwhile:
TTLremains a final safety net.
This does not give perfect cross-system atomicity, but it provides a reliable recovery path.
Cache-Aside vs Write-Through vs Write-Behind
| Strategy | Read behavior | Write behavior | Main benefit | Main cost |
|---|---|---|---|---|
| Cache-aside | App loads on miss | DB then invalidate/update | Simple and flexible | Miss logic + stale windows |
| Read-through | Cache abstraction loads on miss | Separate write policy | Centralized read policy | Can hide expensive misses |
| Write-through | Cache updated during write | DB + cache synchronously | Warm cache after writes | Slower writes + dual-system failures |
| Write-behind | Usually cache/buffer first | DB updated later | High write throughput | Durability, ordering, replay complexity |
When to Use Cache-Aside
Good default when:
database is source of truth
reads dominate
staleness is manageable
application wants direct policy controlThis is often the first caching strategy to consider.
When to Use Write-Through
Useful when:
written data is read immediately
warm-cache guarantee is valuable
extra write latency is acceptableBut define partial-failure behavior clearly.
When to Use Write-Behind
Use only when:
delayed persistence is acceptable
write throughput matters greatly
durable buffering exists
replay and ordering are designedIt is usually not the default for critical CRUD data.
When to Use Read-Through
Useful when:
many services need
the same read caching policyand a shared abstraction improves consistency.
Still keep misses observable.
Consistency Is a Product Decision
Ask:
How stale may users see this value?Possible answers:
0 seconds
5 seconds
1 minute
10 minutesDifferent answers lead to different architecture.
Example:
homepage recommendations
→ maybe 10 minutes
product description
→ maybe 5 minutes
price
→ maybe seconds
permission revocation
→ perhaps no stale read allowedThere is no "best Redis TTL" independent of the business requirement.
Cache Stampede and Database Protection Are Connected
Suppose the database handles:
1,000 queries/secnormal cache miss traffic:
300/secFine.
If a hot key expires:
5,000 misses/secthe database cannot handle it.
Stampede protection therefore is not only:
cache optimizationIt is:
database overload protectionAdmission Control on Cache Misses
In extreme systems, even miss rebuilding may need concurrency limits.
For example:
maximum 50 expensive cache rebuilds
at onceAdditional misses may:
serve stale
wait briefly
or
failinstead of launching unlimited DB queries.
This connects caching with backpressure.
Hot-Key Refresh and Single-Flight Together
For a highly valuable hot key:
proactive refresh
+
stale-while-revalidate
+
single-flight rebuildcan be combined.
For example:
normal:
refresh before expiry
if refresh fails:
continue serving stale briefly
if key disappears:
only one request rebuildsProduction caching often uses several defenses together.
Avoid Cache Penetration
Cache penetration means repeated requests target values that are never cached successfully.
Example:
random product IDsEvery request:
Redis miss
→ DB missNegative caching can help.
Other protections include:
input validation
rate limiting
Bloom filters in some systems
identity limitsdepending on the workload.
Avoid Cache Avalanche
Cache avalanche generally refers to many keys becoming unavailable together.
Possible causes:
many synchronized TTL expirations
Redis node failure
large namespace invalidationResult:
huge miss spike
→ database overloadMitigations:
TTL jitter
stale serving
rebuild limits
Redis high availability
database admission controlStampede vs Avalanche
Easy distinction:
Stampede
→ many requests rebuild
one or a few hot keysAvalanche
→ many keys disappear
around the same timeBoth can overload the database.
Redis Cluster Does Not Automatically Fix Hot Keys
Suppose you have:
10 Redis shardsbut one key receives:
100,000 requests/secThat key still belongs to:
one shardAdding more shards does not automatically split one key.
You may need:
local caching
replicated reads
application-level partitioning
or redesigndepending on strictness and data shape.
Avoid Caching Sensitive Data Without a Plan
Redis may contain:
user profiles
tokens
permissions
business dataConsider:
access controls
encryption in transit
network isolation
TTL
logging
backup policyDo not put secrets in caches simply because Redis is "internal."
Serialization Strategy
Common choices:
JSON
MessagePack
Protobuf
custom binaryJSON is simple and debuggable.
Trade-offs:
larger values
serialization CPU
type restorationFor most web backends, JSON is often sufficient until measurements show otherwise.
Cache Schema Evolution
Suppose application version A caches:
{
"firstName": "Faisal"
}Version B expects:
{
"name": {
"first": "Faisal"
}
}During rolling deployments both versions may be active.
Versioned keys help:
user:42:v1
user:42:v2This prevents incompatible application versions from sharing one serialized shape.
Deployment and Cache Compatibility
A production deployment may temporarily contain:
Old application instances
+
New application instancesIf both read/write the same incompatible cache format:
errors
wrong datamay occur.
Treat cached representation changes like schema changes.
A Good Redis Failure Flow
Suppose:
GET /products/42normally uses Redis.
Redis times out.
Possible safe flow:
Redis timeout
↓
fallback admission control
↓
capacity available?
┌──┴───┐
yes no
↓ ↓
DB stale/503Do not let every request instantly fall through.
Cache Data That Can Be Reconstructed
An important architectural principle:
cache should normally be disposableIf Redis is flushed:
application may slow down
but should recoverIf Redis is flushed and permanent user data disappears:
that was not merely a cacheTreat it accordingly.
Common Caching Mistakes
Mistake 1: Caching Everything
Not every DB read needs Redis.
If PostgreSQL already returns:
2 msadding:
Redis network call
serialization
invalidation logicmay not help.
Mistake 2: Long TTL Without Staleness Analysis
Do not choose:
1 hourjust because it reduces DB traffic.
Ask whether users can safely see data that old.
Mistake 3: No TTL Because "We Invalidate on Write"
Invalidation can fail.
Keep TTL as a safety net.
Mistake 4: No Stampede Protection
A hot key expiry can make the cache useless exactly when traffic is highest.
Mistake 5: Treating Database as Unlimited Fallback
Redis outages can send all traffic to PostgreSQL.
Bound fallback.
Mistake 6: Global Hit Rate Only
A 99% overall hit rate can still hide one endpoint destroying the database.
Mistake 7: Huge Cache Values
Large values create network, memory, and serialization problems.
Mistake 8: Stale Authorization Data
Performance is not worth incorrect security decisions.
Mistake 9: Using Redis as Durable Storage by Accident
If data cannot be rebuilt, define Redis durability requirements explicitly.
Mistake 10: No Cache Versioning
Rolling deployments can deserialize incompatible old values.
Use versioned schemas/keys when representations change.
A Practical Decision Process
Before caching anything, ask:
1. What problem are we solving?For example:
DB latency?
DB CPU?
external API cost?
high read volume?Then:
2. Is this data reused enough
to produce cache hits?Then:
3. How stale may it safely become?Then:
4. What is the source of truth?Then:
5. How is cache populated?Then:
6. What happens after writes?Then:
7. What happens when a hot key expires?Then:
8. What happens when Redis is slow
or unavailable?Then:
9. Can the database survive
the miss traffic?Finally:
10. How will we observe whether
the cache is actually helping?Only then should you implement the Redis calls.
Example Architecture
A mature read path might look like:
Request
↓
small local cache
↓ miss
Redis
↓ miss
single-flight gate
↓
PostgreSQL
↓
serialize
↓
Redis SET
TTL + jitter
↓
returnThe local layer may be skipped if unnecessary.
Example Write Architecture
PATCH product
↓
PostgreSQL transaction
↓
update product
+
write outbox event
↓
commit
↓
responseThen:
outbox worker
↓
invalidate Redis
↓
optionally notify local cachesAnd TTL remains the safety net.
Production Checklist
Before shipping a Redis cache:
□ Define the source of truth.
□ Define acceptable staleness.
□ Choose cache-aside,
read-through,
write-through,
or write-behind deliberately.
□ Decide how cache entries
are populated.
□ Decide what happens after writes.
□ Keep a safety TTL.
□ Add TTL jitter for large
related key groups.
□ Version cache keys
when schemas change.
□ Keep values reasonably small.
□ Avoid huge unbounded lists.
□ Use deterministic namespaces.
□ Avoid unsafe wildcard invalidation.
□ Protect hot-key rebuilds.
□ Use request coalescing
where useful.
□ Consider stale-while-revalidate.
□ Consider proactive refresh
for known hot keys.
□ Use short negative-cache TTLs.
□ Bound negative-cache cardinality.
□ Monitor hot keys.
□ Use local caching only
where extra staleness is safe.
□ Define Redis timeout behavior.
□ Bound database fallback concurrency.
□ Do not assume the DB can absorb
all cache traffic.
□ Define what happens during
Redis outage.
□ Monitor hit and miss rates
by endpoint/namespace.
□ Measure miss/rebuild latency.
□ Monitor Redis memory.
□ Monitor key count.
□ Monitor evictions.
□ Monitor network throughput.
□ Monitor cache errors/timeouts.
□ Measure data age when
freshness matters.
□ Test cache stampedes.
□ Test Redis latency.
□ Test complete Redis outage.
□ Test rolling deployments
with cache schema changes.
□ Keep correctness-sensitive
decisions off stale paths.Conclusion
Caching is not simply:
Request
→ Redis
→ DatabaseThe real architecture includes:
population
invalidation
TTL
staleness
concurrent rebuilds
failure behavior
hot keys
database protection
schema evolution
observabilityThe common strategies make different trade-offs.
Cache-Aside
→ application loads on miss
→ simple and flexibleRead-Through
→ cache abstraction loads on miss
→ centralized read policyWrite-Through
→ update database and cache
during the write
→ warm reads afterwardWrite-Behind
→ persist later
→ faster write path,
but much more durability complexityThe hardest part is usually not:
redis.get()
redis.set()It is deciding what happens when the two copies disagree.
A good cache design therefore starts with:
How stale may this data safely become?
Then asks:
How do we recover when invalidation, Redis, or the database behaves differently from the happy path?
For production systems, also remember:
Cache hit
→ fast path
Cache miss
→ expensive path
Redis outage
→ potentially dangerous pathAll three need to be designed.
The final rule is simple:
A cache should make the system faster without becoming a hidden source of incorrect data or uncontrolled database load.
If you can explain the hit path, miss path, write path, stale-data window, stampede behavior, and Redis-outage behavior for every important cached value, then you have a real caching architecture—not just Redis in front of PostgreSQL.
When Redis appears healthy but PostgreSQL is still overloaded, use the cache-failure diagnostic guide to trace misses, synchronized expiry, hot keys, and database work per miss. If the core question is how much stale data users may observe, the consistency-models guide provides the broader contract.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.