Short answers to what readers ask most about this topic.
01Why do read replicas return stale data?
A replica receives and replays the primary's write-ahead log asynchronously. The primary confirms a commit before the replica has applied it, so for a short window the replica still holds the old rows. That gap is replication lag.
02What is read-your-writes consistency?
It is the guarantee that after a process writes, its later reads reflect that write. With read replicas it is not automatic, so you provide it by sending a user's reads to the primary for a short window after they write, or by waiting until a replica has replayed the write's LSN.
03How do I measure replication lag in Postgres?
On the primary, query pg_stat_replication for replay_lag, or use pg_wal_lsn_diff between pg_current_wal_lsn() and replay_lsn for bytes. On a replica, now() minus pg_last_xact_replay_timestamp() gives an apparent lag. That last figure grows on an idle primary even when the replica is caught up.
04Should I read from the primary after every write?
Only for the user who just wrote, and only for a short window or until a replica reaches their write position. Sending all reads to the primary defeats the purpose of replicas. A short per-user or per-record pin keeps most reads on replicas.
It makes a commit wait until the synchronous standby has applied the transaction, so that standby can serve it. The cost is much larger commit delays, and it only applies when synchronous_standby_names is set. Use it for a few critical transactions, not as a global default.
Read replicas return stale data because they apply the primary's write-ahead log asynchronously, so a replica always trails the primary by some lag. Handle it by measuring lag with pg_stat_replication, sending reads that follow a user's own write to the primary for a short window, and keeping each session on one replica.
Picture a POS screen. The cashier saves a wash ticket, the app redirects to the ticket page, and the page answers that the ticket does not exist. Two seconds later a refresh shows it. Nothing was lost; the write went to the primary and the read went to a replica that had not caught up yet.
This post is about that gap: why it exists, what it looks like to users, how to measure it in Postgres, and three fixes in order of cost. The numbers in the worked examples are illustrative arithmetic, and every behaviour claim about Postgres comes from the official documentation listed at the end.
Why do read replicas return stale data?
A streaming replica is a second server that receives the primary's write-ahead log (WAL) and replays it. By default the primary does not wait for that to finish before it tells the client the commit succeeded, so there is always a window where the primary has the row and the replica does not. Postgres reports that window in three stages, one per column of pg_stat_replication.
Column
What it measures
Matching synchronous_commit level
write_lag
WAL received and written by the standby, not yet flushed or applied
remote_write
flush_lag
WAL written and flushed to durable storage on the standby, not yet applied
on
replay_lag
WAL applied, so the change is visible to queries on the standby
remote_apply
Only the last stage matters to a user, because a query cannot see a change until it is replayed. A replica can have the WAL on disk and still return the old row. Replay can also be delayed on purpose by the standby itself: with hot standby, a query that conflicts with incoming WAL can hold replay back for up to max_standby_streaming_delay, which defaults to 30 seconds.
What does replication lag look like to users?
Lag shows up as two distinct bugs that deserve different fixes. A user who writes and then reads can miss their own change. A user who reads twice can watch data go backwards, because the two reads were served by replicas at different positions.
Symptom
Guarantee violated
Fix
Saved a record, the next page says it does not exist
Read-your-writes
Route the read to the primary after a write
A row appears on one refresh and vanishes on the next
Monotonic reads
Keep each session on one replica
A dashboard total is a few seconds behind
None, this is plain eventual consistency
Accept it, and show an as-of time
A worked example. A user commits at t = 0. The app redirects and the next request arrives at t = 120 ms. If the replica's replay lag is 800 ms, that read runs 800 minus 120 = 680 ms too early and misses the row. Now add a second replica with a lag of 50 ms. A request at t = 500 ms hits the fast replica and sees the row; one at t = 600 ms hits the slow replica, which will not apply the commit until t = 800 ms, and the row disappears. Read-your-writes means a process sees its own earlier writes; monotonic reads means a later read never shows an older state than an earlier one. Both are defined per process, which is why a per-session fix works.
How do you measure replication lag in Postgres?
Measure from both sides. The primary knows how far each standby has got, and the replica knows how old its newest replayed commit is. The first query below reports lag in bytes and as time per standby; the second is what you can run when you only have access to the replica; the third checks whether a replica has replayed a specific position.
-- On the PRIMARY: one row per connected standby.
SELECT application_name,
state,
pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS replay_lag_bytes,
write_lag,
flush_lag,
replay_lag -- what a reader on that standby actually feels
FROM pg_stat_replication;
-- On the REPLICA: how old is the newest commit I can see?
SELECT now() - pg_last_xact_replay_timestamp() AS apparent_lag;
-- On the REPLICA: has it replayed at least the LSN I was handed after a write?
SELECT pg_wal_lsn_diff(pg_last_wal_replay_lsn(), '0/3000148') >= 0 AS caught_up;
pg_wal_lsn_diff subtracts two WAL locations and returns bytes, which the documentation points to for working out lag. pg_last_xact_replay_timestamp returns the time the last replayed transaction was committed on the primary, so subtracting it from now() gives an apparent lag. The third query is the building block for an exact after-write check: capture pg_current_wal_lsn() on the primary after the commit, and later ask a replica whether its replayed position has reached it.
now() minus pg_last_xact_replay_timestamp() is only honest while the primary is writing. On a quiet primary the last replayed commit simply gets older, so the number grows even though the replica is fully caught up. Likewise pg_stat_replication shows the time to replay the most recent WAL rather than zero when caught up, then NULL after a while. Treat NULL as no recent measurement, not as zero lag.
How do you route reads to the primary after a write?
The cheapest fix that works is a short pin. After any write made for a user, remember that fact for a few seconds, and while it is set send that user's reads to the primary. The NestJS service below does it with a Redis key that expires on its own, and picks a replica for everyone else. The window is a design choice: 5 seconds here is an alert threshold of 2 seconds, plus a 1 second lag poll, plus margin.
import { Injectable } from "@nestjs/common";
import { Pool } from "pg";
import Redis from "ioredis";
// Wrong: send every SELECT to the replica and hope the lag is small.
// Right: a user who just wrote is pinned to the primary for a short window.
const PIN_SECONDS = 5; // alert threshold 2 s + poll interval 1 s + margin
@Injectable()
export class DbRouter {
constructor(
private readonly primary: Pool,
private readonly replicas: Pool[],
private readonly redis: Redis,
) {}
// Call this after any committed write made on behalf of a user.
async markWritten(userId: string): Promise<void> {
await this.redis.set("ryw:" + userId, "1", "EX", PIN_SECONDS);
}
async poolFor(userId: string, sessionId: string): Promise<Pool> {
if (await this.redis.exists("ryw:" + userId)) return this.primary;
return this.pickReplica(sessionId);
}
// Same session, same replica: this is what keeps reads monotonic.
private pickReplica(sessionId: string): Pool {
let h = 0;
for (const c of sessionId) h = (h * 31 + c.charCodeAt(0)) >>> 0;
return this.replicas[h % this.replicas.length];
}
}
The pin trades primary load for correctness, and only for users who just wrote, so the primary absorbs a small slice of reads rather than all of them. If a pin window ever proves too short, the exact alternative is the LSN check from the previous section: store the LSN returned after the commit instead of a boolean, and read from a replica only when its replayed position has reached it. That removes the guess about duration at the cost of one extra query per read.
Pin by the thing the user wrote, not the whole user, when you can. If a cashier edits one ticket, pinning reads of that ticket id for a few seconds protects the screen that needs it and leaves the rest of the dashboard on replicas.
How do you keep reads monotonic across replicas?
Monotonic reads fail when consecutive requests from one user land on replicas at different positions. The fix is to make the choice of replica stable per user instead of random. Three common ways, from simplest:
Hash the session or user id to a replica, as pickReplica does above. No state, and it survives restarts of the app.
Sticky routing at the load balancer, which gives the same result but moves the logic out of the application.
Pass an LSN token with each request and only read from a replica whose replay position has reached it. This is the strongest option and the most work.
Stability has a cost: if the replica a session is stuck to dies, that session moves to another one and may briefly see older data once. That is usually an acceptable trade, and it is the same compromise Jepsen describes for these guarantees, which hold only while a client keeps talking to the same server.
Can synchronous_commit remote_apply fix stale reads?
Yes, with a price. With remote_apply a commit waits until the synchronous standby has applied the transaction, so it is visible to queries there. That needs synchronous_standby_names to be non-empty; it has no effect otherwise. The documentation warns it causes much larger commit delays than the other levels because it waits for replay.
# postgresql.conf on the primary
synchronous_standby_names = 'ANY 1 (replica_a, replica_b)'
# Per transaction, only where a stale read is unacceptable:
# SET LOCAL synchronous_commit = 'remote_apply';
# Everything else keeps the default and does not wait for replay.
That is why I treat it as a scalpel. Setting it for a handful of transactions where a stale read is unacceptable is reasonable; setting it globally turns every write into a wait on the slowest synchronous replica, and the primary stalls if that standby is unreachable. With ANY 1 over two replicas, a commit needs only one of them, but a read can still hit the other, so remote_apply alone does not guarantee a replica read sees the write.
When should reads fall back to the primary?
Use this checklist to decide, in order, where a read goes. It encodes the fixes above so the rule lives in one place instead of in each endpoint.
Did this user write in the last pin window? Read from the primary.
Is the read part of a transaction that also writes? Use the primary connection.
Is the chosen replica's replay_lag above your threshold, or its state not streaming? Fall back to the primary.
Otherwise pick the replica by hashing the session id so reads stay monotonic.
For anything that may be seconds old by design, such as reports, show an as-of timestamp rather than hiding the lag.
Step 3 is the one that protects the primary from a replica outage: if every replica is lagging, everything falls back at once, so size the primary for that worst case or shed the heaviest reads first.
Replication lag is not a bug to remove but a property to route around. Measure replay lag, pin a user to the primary right after they write, keep each session on one replica, and spend remote_apply only where a stale read costs more than a slower commit.