Short answers to what readers ask most about this topic.
01What is the best way to invalidate a cache?
There is no single best way; match the method to how stale a read may be. Put a TTL on every key as a safety net, delete or version the key after the database commit when data must be fresh, and add jitter plus single-flight to avoid stampedes. Most production systems use all three together.
02Should I use TTL or event-based cache invalidation?
Use a TTL when slightly stale data is acceptable, because worst-case staleness equals the TTL and nothing else can fail. Use event-based invalidation when a stale read costs money, such as a price or stock quantity, but keep a TTL as the backstop in case a purge is lost.
03What is a cache stampede and how do I prevent it?
A cache stampede happens when a popular key expires and many requests miss at once, all running the same expensive query. Prevent it with single-flight so concurrent callers share one rebuild, a lock such as SET NX PX in Redis, jitter on TTLs, or probabilistic early expiration.
04Should I delete the cache key or update it when data changes?
Deleting after the database commit is the safer default, because the next read rebuilds from the source of truth and you avoid writing a value computed from a half-finished update. Updating in place only makes sense when you can write the exact new value inside the same step. Never purge before the commit.
05How do I delete many Redis keys by pattern without blocking?
Avoid KEYS, which walks the whole keyspace in a single command. Use redis-cli --scan with a pattern and pipe the results to UNLINK in batches, since UNLINK reclaims memory in another thread. Better still, put a version counter in the key so one INCR retires the whole group.
Cache Invalidation Strategies: TTL, Events and Versioned Keys
How to invalidate a cache correctly: why reads go stale, how TTL, event-driven purge and versioned keys differ, and how jitter and single-flight stop stampedes.
Invalidate a cache by matching the method to how stale a read may be. Keep a TTL as the safety net, delete or version the key after the database commit when reads must be fresh, and add jitter plus single-flight so one expiry causes one rebuild, not a stampede. No single technique covers all three failure modes.
Picture a carwash ERP where the owner raises the price of a wash package from 10,000 to 12,000 rupiah. The database says 12,000, the cashier tablet still shows 10,000, and nobody can say for how long. That gap between the source of truth and the copy is what cache invalidation is about.
I build ERP and POS systems on NestJS, Postgres and Redis, so this post uses that stack. It answers the question people actually search: how do you invalidate a cache correctly with TTLs, events or versioned keys? Every number below is worked out in the text, and every Redis behaviour is traced to the official documentation listed at the end.
Why does a cache return stale data after the database changed?
A cache is a second copy, and two copies disagree whenever one changes without the other. The obvious fix, delete the key when you write, still leaves a window. In the cache-aside pattern a reader that started before the write can finish after the delete and put the old value back. Here is that sequence with a 60 second TTL.
t0 reader A: GET product:42 -> miss
t1 reader A: SELECT price ... -> 10000 (old row)
t2 writer B: UPDATE price = 12000; COMMIT
t3 writer B: DEL product:42 -> key did not exist yet, nothing to delete
t4 reader A: SET product:42 10000 EX 60 -> the old price is cached again
# Result: every reader sees 10000 until t4 + 60 s, although the database
# says 12000 and the invalidation "worked".
So a purge on write does not guarantee freshness by itself. It shrinks the stale window from the whole TTL to the rare interleaving above, and only if the purge happens after the commit. The TTL on every key is what bounds the damage when the interleaving does occur, which is why event-driven invalidation and TTLs are partners, not rivals.
Never treat delete-on-write as proof that readers are fresh. A reader that loaded the row before your commit can re-cache it after your delete, and the stale value then lives for the full TTL. Keep a TTL on every key, even keys you purge explicitly.
How does a TTL work and how long should it be?
A TTL is the simplest invalidation there is: set the key with an expiry, for example SET key value EX 60 in Redis, and the entry disappears on its own. The worst-case staleness equals the TTL. If updates land at random moments, the average age of a wrong value is about half of it, so 30 seconds on a 60 second TTL. Choose the TTL from the staleness the business tolerates, not from what feels tidy.
The cost is rebuild work. A key read constantly is rebuilt once per TTL: 3600 divided by 60 gives 60 rebuilds an hour at a 60 second TTL, and 3600 divided by 5 gives 720 an hour at 5 seconds, twelve times the database load for the same key. Now add synchronised expiry. If 10,000 keys are warmed together with a 300 second TTL they all expire in the same second. Adding plus or minus 10 percent jitter spreads them over 270 to 330 seconds, a 60 second window, so the burst becomes roughly 10,000 divided by 60, about 167 rebuilds per second.
What is a cache stampede and how do single-flight and jitter prevent it?
A stampede, also called a thundering herd or dogpile, happens when a popular key expires and many requests miss at once, each running the same expensive query. Take a key read 200 times a second whose rebuild takes 400 ms. During the rebuild 200 times 0.4 gives 80 requests arrive, so 80 identical queries hit Postgres instead of 1. If the pool holds 10 connections, which is an assumption for this example, 70 of them queue and the slow rebuild gets slower. Wikipedia lists three mitigations: locking, external recomputation and probabilistic early expiration. The code below combines the first with jitter.
import Redis from "ioredis";
import { randomUUID } from "node:crypto";
const redis = new Redis();
const inflight = new Map<string, Promise<string>>();
const BASE_TTL = 300; // seconds
const JITTER = 0.1; // +/- 10 percent -> 270..330 s
const ttlWithJitter = () =>
Math.round(BASE_TTL * (1 - JITTER + Math.random() * 2 * JITTER));
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
// Delete the lock only if it is still ours (the pattern from the Redis SET docs).
const UNLOCK = `if redis.call("get",KEYS[1]) == ARGV[1]
then return redis.call("del",KEYS[1]) else return 0 end`;
export async function getCached(
key: string,
load: () => Promise<string>,
): Promise<string> {
const hit = await redis.get(key);
if (hit !== null) return hit;
// Layer 1: single-flight inside this process. 80 callers, one promise.
const pending = inflight.get(key);
if (pending) return pending;
const run = (async () => {
// Layer 2: one rebuilder across every process that shares this Redis.
const lockKey = "lock:" + key;
const token = randomUUID();
const got = await redis.set(lockKey, token, "PX", 5000, "NX");
if (got !== "OK") {
await sleep(50); // another process is rebuilding; look again shortly
const again = await redis.get(key);
if (again !== null) return again;
// Still empty: load anyway. A slow rebuilder must not fail every reader.
}
try {
const value = await load();
await redis.set(key, value, "EX", ttlWithJitter());
return value;
} finally {
if (got === "OK") await redis.eval(UNLOCK, 1, lockKey, token);
}
})().finally(() => inflight.delete(key));
inflight.set(key, run);
return run;
}
Two layers do different jobs. The in-process map is single-flight: concurrent callers in one Node process share one promise, so 80 callers become one. The Redis lock, SET with NX and a PX expiry, picks one rebuilder across all processes. The lock carries a random token and is released by a script that deletes it only when the token still matches, the pattern shown in the Redis SET documentation, so a slow client never deletes a lock another client has since taken. Redis documents the plain SET NX EX lock as discouraged in favour of Redlock when you need stronger guarantees, and for cache rebuilds a lost lock only costs a duplicate query.
The alternative is to avoid the miss altogether. Probabilistic early expiration, called XFetch on Wikipedia, makes each reader recompute early when now minus delta times beta times ln(U) is at or past the expiry, where delta is the last rebuild time, U is a random number between 0 and 1, and beta defaults to 1. With a 0.4 second rebuild, a reader refreshes on average about 0.4 seconds before expiry, since the mean of minus ln(U) is 1. No lock is needed, but you must store delta next to the value.
When should you invalidate with events instead of waiting for a TTL?
Use event-driven invalidation when a stale read has a cost a TTL cannot bound: a price, a stock quantity, a permission. The rule is about ordering. Write to the database, commit, and only then remove the key. Purging before the commit lets a reader refill the old row in the gap, and the wrong value then survives until its TTL.
// Wrong: purge first, write second. A reader refills the OLD row in between.
await redis.del("product:" + id);
await db.query("UPDATE products SET price = $1 WHERE id = $2", [price, id]);
// Right: the purge is the LAST step, after the commit is durable.
await db.query("BEGIN");
await db.query("UPDATE products SET price = $1 WHERE id = $2", [price, id]);
await db.query("COMMIT");
// UNLINK frees memory in another thread; DEL does it on the main one.
await redis.unlink("product:" + id);
// Optional: tell other app instances to drop their in-process copy.
// Pub/Sub is fire and forget, so this is a hint. The TTL stays as the backstop.
await redis.publish("cache:purge", "product:" + id);
Two details matter. UNLINK removes the key like DEL but reclaims memory in another thread, so a large value does not block the server. And Redis Pub/Sub is fire and forget: the keyspace notifications page states that events published while a subscriber is disconnected are lost. A purge broadcast is therefore an optimisation for in-process copies, never the only line of defence. If a missed purge is unacceptable, write the purge to an outbox table in the same transaction and let a worker retry it.
In an ERP the write path is usually a few well-known service methods. Put the invalidation in the same method that commits, behind one helper, instead of scattering DEL calls through controllers. One place to audit beats forty.
What are versioned keys and when are they better than deleting?
A versioned key puts a counter in the key name: product:42:v7. To invalidate, increment the counter. The old entry is never read again and simply ages out by TTL, so there is nothing to find and delete. It also closes the race from the first section, provided the reader captures the version before it reads the database and the writer increments after the commit. A reader holding version 7 that loaded an old row writes to product:42:v7, a key nobody will ask for once the counter reads 8.
// The version counter is the only thing a writer touches.
const verKey = (id: string) => "ver:product:" + id;
export async function readProduct(id: string) {
// Capture the version BEFORE the database read, never after.
const v = (await redis.get(verKey(id))) ?? "0";
const key = "product:" + id + ":v" + v;
const hit = await redis.get(key);
if (hit !== null) return JSON.parse(hit);
const row = await loadProductFromPostgres(id);
await redis.set(key, JSON.stringify(row), "EX", 3600);
return row;
}
export async function onProductCommitted(id: string) {
// Run after COMMIT. Old keys are unreachable at once and expire by TTL.
await redis.incr(verKey(id));
}
The price is one extra round trip per read to fetch the version, and dead keys occupy memory until their TTL ends, so set a TTL and a maxmemory policy. In return one INCR on a namespace such as ver:products retires every list and detail key built on it, which is how you invalidate groups of keys without scanning for them. For a hot path, keep the version in a short-lived in-process variable and accept a few seconds of lag.
How do you invalidate many Redis keys without blocking the server?
Resist the temptation to purge by pattern. KEYS walks the whole keyspace in a single command, which is the wrong thing to run on a Redis that also serves your sessions and rate limits. SCAN returns keys in batches, and UNLINK deletes without blocking, so the safe version streams matches 100 at a time.
# Wrong: KEYS walks the whole keyspace in one command, then DEL frees on the main thread.
redis-cli KEYS 'product:*' | xargs redis-cli DEL
# Right: SCAN in batches, UNLINK 100 keys at a time, reclaim memory off-thread.
redis-cli --scan --pattern 'product:*' | xargs -n 100 redis-cli UNLINK
# Better still: no scan at all. One INCR retires every key under a namespace.
redis-cli INCR ver:products
Even the safe version is a batch job, not an invalidation strategy. Between the scan and the last UNLINK, readers can refill keys you already purged, and on a Redis Cluster the keys live on different nodes. If you regularly need to drop a group, that is the signal to design a version counter into the key from the start.
Which cache invalidation strategy should you choose?
The five options below are not alternatives to rank. Most systems combine a TTL, one freshness mechanism and a stampede guard. The table states what each one actually guarantees.
Strategy
Worst-case staleness
Main failure mode
Best fit
TTL only
The full TTL
Synchronised expiry and stampedes
Reference data that changes rarely, such as a service catalogue
TTL with jitter and single-flight
The full TTL
Lock expires before the rebuild finishes, causing a duplicate query
Hot keys with an expensive rebuild
Delete after commit
Near zero, but the full TTL if a reader re-caches an old row
The delete is lost or runs before the commit
Prices, stock and permissions where staleness is costly
Versioned keys
Near zero, with no refill race
Extra read per request and dead keys using memory until TTL
Groups of keys that must be retired together
Stale-while-revalidate
TTL plus the stale window
Users briefly see old data by design
Dashboards and feeds where speed beats exactness
Stale-while-revalidate is standardised for HTTP caches in RFC 5861: Cache-Control max-age=600, stale-while-revalidate=30 means a response is fresh for 600 seconds and may be served stale for 30 more while a background refresh runs. The same idea works in application code. Then walk this checklist for each key.
Write down the staleness the business accepts, in seconds, and set the TTL to that figure at most.
Add 10 percent jitter to the TTL of any key that is warmed in bulk.
If staleness above a few seconds costs money, delete or version the key after the commit, and keep the TTL as the backstop.
Put single-flight or a Redis lock on any key whose rebuild takes longer than a few tens of milliseconds and is read often.
Cache invalidation is hard because three problems hide inside one name: how stale a read may be, how many rebuilds one expiry causes, and how to retire a group of keys. Solve each on purpose. A TTL bounds staleness, delete-after-commit or a version counter makes it fresh, and jitter with single-flight keeps one expiry from becoming eighty queries.