You Added Read Replicas to Scale Reads… So Why Are Users Seeing Old Data After an Update?
Why replication lag causes stale reads after writes, and how primary routing, consistency windows, session stickiness, version checks, and LSN-based techniques restore read-after-write behavior.
7 min read
The Update Succeeded, but the Next Page Shows the Old Value
A user changes their shipping address.
PATCH /profile/address -> 200 OK
GET /profile -> old addressThey refresh again two seconds later and see the new address.
The write was not lost. The first read went to a replica that had not replayed the write yet.
write
Client ------------------------> Primary
|
| replication stream
v
Replica
^
Client ---------------------------|
immediate readRead replicas increase capacity by accepting weaker freshness unless the application adds a consistency strategy.
Why Replicas Lag
PostgreSQL commonly streams WAL records from the primary. A replica must receive and replay those records.
Primary commit
|
v
WAL generated
|
v
WAL sent over network
|
v
Replica receives WAL
|
v
Replica replays changeLag can increase because of:
- Network delay
- Heavy write volume
- Slow replica storage
- Replica CPU saturation
- Long-running queries on the replica
- Recovery conflicts
- Maintenance or failover events
Typical lag may be small, but correctness cannot rely on “usually under 100 ms” if the product requires an immediate consistent read.
Eventual Consistency Is Not Automatically Wrong
Many reads tolerate staleness:
- Product recommendations
- Analytics dashboards
- Public counters
- Historical reports
- Search results
Other reads often require fresher behavior:
- “Show the profile I just saved”
- Permission changes
- Order confirmation
- Account balance after transfer
- Inventory reservation result
Define consistency per use case instead of forcing every read to the primary or every read to a replica.
Solution 1: Return the Updated Representation
Avoid an immediate second read when the write already knows the result.
UPDATE user_profiles
SET address = $1,
version = version + 1
WHERE user_id = $2
RETURNING user_id, address, version, updated_at;Return that row in the write response. The client can update its local state without reading a replica immediately.
This does not solve later reads from another device, but it removes a common self-inflicted stale read.
Solution 2: Read From the Primary After a Write
For a bounded consistency window:
user writes at 10:00:00
reads by that user until 10:00:05 -> primary
later reads -> replicaThe API can set a short-lived consistency cookie or session flag:
res.cookie("primary-read-until", Date.now() + 5_000, {
httpOnly: true,
sameSite: "lax",
secure: true,
});Routing logic chooses the primary while the window is active.
Benefits:
- Simple mental model
- Strong read-after-write behavior for the same user
- Replica capacity remains useful for other reads
Costs:
- Primary receives additional reads
- A fixed window is only an estimate
- The signal must follow the user across app instances
Solution 3: Route Critical Reads to the Primary
Some endpoints should always use primary data:
GET /checkout/availability -> primary
GET /account/balance -> primary
GET /catalog/search -> replica
GET /analytics/summary -> replicaThis is explicit and reliable, but it reduces the amount of read traffic replicas can absorb.
Keep primary and replica data-access APIs visibly different so a later refactor does not accidentally move a critical read.
Solution 4: Read Until a Required Version
If rows have monotonically increasing versions, the write response can return version 18.
The next read requires at least version 18:
required version: 18
replica returns: 17 -> not fresh enough
fallback primary: 18 -> returnThis works well for entity reads where the application can compare a version.
const replicaProfile = await replicas.profile.findUnique({ where: { userId } });
if (!replicaProfile || replicaProfile.version < requiredVersion) {
return primary.profile.findUniqueOrThrow({ where: { userId } });
}
return replicaProfile;The required version can travel in a short-lived token or client request header.
Solution 5: Wait for a WAL Position
Advanced PostgreSQL designs can record the primary’s WAL position after a write and ensure a replica has replayed through that position before serving the dependent read.
Conceptually:
write commits at WAL position X
|
v
read asks for data at least through X
|
+--> replica caught up: read replica
|
+--> timeout: fallback primary or fail explicitlyThis is more precise than a fixed time window but creates infrastructure and routing complexity. Use it when consistency requirements justify it.
What About Synchronous Replication?
Synchronous replication can require one or more standbys to acknowledge WAL before commit returns.
This reduces the risk of data loss during failover and can improve replica freshness guarantees, depending on configuration. It also adds write latency and can reduce write availability if required standbys are unavailable.
It does not mean every load-balanced read automatically goes to a replica that has applied the latest change. Understand the configured acknowledgment level and replay behavior.
Cache Can Reintroduce Staleness
Even if the database read goes to the primary, Redis may still return an older value.
write primary
|
v
read cache -> stale valueInvalidate or version relevant cache entries as part of the write strategy. The guide on Redis caches that still overload the database also covers invalidation trade-offs.
Consistency is end-to-end. Solving replica lag while ignoring HTTP, CDN, application, and Redis caches gives users the same symptom through another layer.
Measure Replica Lag
On a primary, inspect replay state for replicas:
SELECT
application_name,
client_addr,
state,
sent_lsn,
write_lsn,
flush_lsn,
replay_lsn,
replay_lag
FROM pg_stat_replication;On a standby:
SELECT
now() - pg_last_xact_replay_timestamp() AS replay_delay;Interpret time-based lag carefully when the system is idle. Also monitor WAL distance, replay throughput, replica CPU, storage latency, and long-running queries.
Useful application metrics:
replica_read_total
primary_fallback_total
read_after_write_primary_total
replica_version_too_old_total
replica_lag_bytes
replica_replay_delay_secondsDiagnose the User Report
Correlate one request flow:
write request ID
commit timestamp
primary WAL position or entity version
next read request ID
database target: primary or replica
returned row version
cache outcome
replica lag at that momentWithout recording which database served the read, “the database returned old data” is difficult to prove.
Trade-offs
| Strategy | Consistency | Cost |
|---|---|---|
| Always read replica | Eventual | Stale reads possible |
| Return write result | Immediate for current client | Does not govern future reads |
| Timed primary stickiness | Stronger read-after-write | More primary load; heuristic window |
| Critical routes on primary | Strong for selected workflows | Less read scaling |
| Version-aware fallback | Precise per entity | Version propagation complexity |
| WAL-position routing | Precise infrastructure signal | Operational and routing complexity |
| Synchronous replication | Stronger durability/freshness | Write latency and availability cost |
Production Best Practices
- Classify which reads require read-after-write consistency.
- Return updated state directly from write endpoints.
- Route critical reads to the primary.
- Use a bounded primary-read window for user sessions when appropriate.
- Carry entity versions for precise fallback decisions.
- Measure lag in time and WAL distance.
- Record which database served each traced request.
- Include caches and CDNs in consistency analysis.
- Test lag during write bursts and replica recovery.
- Protect primary capacity when adding fallback traffic.
Conclusion
Users see old data because the write and the next read reached different points in the replication timeline.
The solution depends on the required guarantee:
eventual read -> replica
read your own write -> return write result or temporary primary routing
critical current read -> primary
precise freshness -> version or WAL-position-aware routingRead replicas scale reads, but the application still owns the decision about how stale each read is allowed to be.
The consistency-models guide places read-your-writes alongside strong, eventual, monotonic, and causal consistency guarantees.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.