Short answers to what readers ask most about this topic.
01How does a CDN work?
A CDN keeps copies of your responses on edge servers, called points of presence, in many locations. A visitor is routed to a nearby one by DNS or anycast. On a hit it answers from its cache, on a miss it fetches from your origin, stores the result and serves it to later visitors.
02What should you cache at the edge?
Cache responses that are identical for every visitor and can be slightly old: fingerprinted JS and CSS, images, fonts, public HTML and public API GET responses. Use long lifetimes for assets whose filename contains a content hash, and short ones for HTML. Start from a conservative s-maxage and tune it.
03What is the difference between max-age and s-maxage?
max-age sets the freshness lifetime for every cache, including the browser. s-maxage applies only to shared caches such as a CDN and overrides max-age there. Together they let the edge keep a copy for minutes while the browser revalidates on every visit.
04What does stale-while-revalidate do?
It lets a cache serve a stale copy immediately while it refreshes the object in the background, for a window you specify. With s-maxage=300 and stale-while-revalidate=60, an object aged 340 seconds is still served instantly while the cache fetches a new copy. Visitors never wait for the origin during that window.
05What should you never cache on a CDN?
Never cache responses that depend on a session cookie or Authorization header, responses that set cookies, non-GET requests, or data that must be current such as live stock or payment status. Mark these private, no-store. Caching them risks one visitor receiving another visitor's data.
How a CDN Works: What to Cache at the Edge and What Not To
A concept explainer on points of presence, anycast routing, cache keys, Cache-Control headers, purging, and the responses that must never reach a shared cache.
A CDN is a network of edge servers, called points of presence, that keep copies of your responses close to visitors. Anycast or DNS routing sends each request to a nearby edge, and the cache key plus Cache-Control headers decide what gets reused. Cache static assets and public pages; never cache personalised, authenticated or state-changing responses.
The question comes up the first time a site that was fast in your own city is slow for someone on another continent: how does a CDN actually work, and which of your responses should it keep? The tutorials answer with a vendor dashboard. The mechanism underneath is plain HTTP caching with a geographic twist.
This post is the concept explainer, not a setup guide. It covers points of presence, how a request finds one, what a cache key is, the exact Cache-Control directives that matter at the edge, what to cache, what to never cache, and how to purge. Every header and its behaviour is taken from MDN and the IETF RFCs listed at the end, and every number is arithmetic you can check.
What is a CDN and how does it work?
A content delivery network is a geographically distributed set of servers that sit between visitors and your origin. Each location is called a point of presence, or PoP, and it runs a shared cache. The first visitor to ask a PoP for an object causes a miss, so the PoP fetches it from your origin and stores it. Later visitors near that PoP get a hit, served without touching the origin.
Two things follow. First, every PoP has its own cache, so a cold object can cost one origin fetch per PoP per lifetime. With 20 PoPs and a 300 second lifetime, that is at most 20 origin fetches every 5 minutes for one object, which is 240 per hour, no matter how many visitors there are. Second, the whole value of the CDN is the hit ratio: at 200 requests per second for an image, a 90 percent hit ratio sends 20 requests per second to the origin, and 99 percent sends 2.
How does a request find the nearest edge?
There are two common routing tricks, and many networks combine them. With DNS-based routing, the authoritative DNS answers the same hostname with different IP addresses depending on where the resolver appears to be. With anycast, many PoPs announce the same IP address and the internet's routing protocol delivers each packet to a topologically close one, as described in the Wikipedia article on anycast.
The word nearest deserves care. Routing optimises for network distance, not kilometres, so a visitor can reach a PoP that is not the closest on a map. For you this changes nothing to configure, but it explains why two users in one city can see different edges, and why a cache hit on one machine tells you nothing about another.
What is a cache key and why does it cause misses?
The cache key is what the edge uses to decide whether two requests want the same object. By default it is built from the URL, and the Vary response header adds request headers to it. Anything that makes keys differ for identical content splits the cache and lowers the hit ratio. The block below shows the usual culprits.
# Default cache key, roughly: scheme + host + path + query string
https://shop.example.com/api/products?page=2&sort=price
https://shop.example.com/api/products?sort=price&page=2 # a DIFFERENT key, same data
# Wrong: tracking parameters in the key. Every campaign link is a miss.
https://shop.example.com/product/42?utm_source=newsletter&fbclid=abc123
# Right: ignore tracking parameters, sort the rest, so one object = one key.
https://shop.example.com/product/42
# Vary adds request headers to the key. This one doubles the cache:
Vary: Accept-Encoding # fine: a handful of values (gzip, br, identity)
Vary: Cookie # fatal: one copy per visitor, hit ratio near zero
The fix is to make one object map to one key: normalise query parameter order, strip tracking parameters where your CDN lets you, and keep Vary to headers with a handful of values such as Accept-Encoding. Check what your own provider includes in its default key before assuming any of this, since the defaults differ between vendors.
Which Cache-Control headers control edge caching?
Cache-Control is the lever, and RFC 9111 defines how caches must read it. Two directives matter most for a CDN. max-age is the freshness lifetime for every cache, and s-maxage overrides it for shared caches such as a CDN, so you can tell browsers one thing and the edge another. The example below keeps the edge copy for five minutes while the browser always asks again.
// NestJS: a public catalogue endpoint the edge may keep, browsers may not.
@Get("products")
@Header(
"Cache-Control",
"public, max-age=0, s-maxage=300, stale-while-revalidate=60, stale-if-error=86400",
)
list() {
return this.products.findPublished();
}
// Fingerprinted build asset: the filename changes when the content changes.
// Cache-Control: public, max-age=31536000, immutable
// Per-user response: shared caches must not store it.
// Cache-Control: private, no-store
stale-while-revalidate and stale-if-error come from RFC 5861. Worked through with the header above: the object is fresh until its Age reaches 300 seconds. From 300 to 360 seconds the cache may serve the stale copy immediately while it refreshes in the background, so no visitor waits. Past 360 seconds it must go to the origin. If the origin is failing, stale-if-error=86400 lets it keep serving the old copy for up to a day. The Age header, defined in RFC 9111, tells you how old the copy you received is.
RFC 9213 adds a header aimed only at CDNs, CDN-Cache-Control, so the edge policy can differ from the browser policy without overloading s-maxage. Support varies by vendor, so confirm it before relying on it. For fingerprinted assets, MDN documents the immutable directive, which tells a browser not to revalidate during the lifetime.
Set max-age=0 with a positive s-maxage on public HTML. The edge absorbs the traffic, but a browser still revalidates, so a purge or a new deploy reaches users on their next request instead of when a long browser lifetime expires.
What should you cache at the edge?
Cache anything that is identical for every visitor and tolerates being a little old. The table sorts common responses by that test, with a header to start from. Treat the lifetimes as starting points to tune, not recommendations measured against your traffic.
The long lifetimes are safe only because the filename changes with the content. That is why the most effective edge optimisation is boring: put a content hash in the URL and cache it for a year.
What should you never cache at the edge?
The failures that matter are not slow pages, they are one person seeing another person's data. In ERP and POS work the pages that show stock, a till session or an invoice are exactly where this would hurt. Keep these out of any shared cache.
Responses that depend on a session cookie or an Authorization header, since a shared cache would store one user's copy and hand it to the next visitor.
Anything carrying Set-Cookie, because caching it can distribute one visitor's session to others.
Non-GET requests. POST, PUT and DELETE change state and must reach the origin every time.
Responses that must be current, such as live stock counts, payment status and anything behind a login.
Sending Vary: Cookie does not make a personalised page safe to cache. It makes one copy per cookie value, which kills the hit ratio while still storing private data on a shared system. Mark personalised responses private, no-store instead.
How do you purge or invalidate cached content?
There are three strategies, in order of preference. Version the URL so nothing needs purging. Use a short lifetime with stale-while-revalidate for content that changes on its own schedule. Purge explicitly only for what you cannot version.
# Strategy 1: version the URL. Nothing to purge, ever.
/assets/app.3f9c2b1e.js # new build = new name = new cache key
# Strategy 2: short TTL plus revalidation for content that changes.
Cache-Control: public, s-maxage=60, stale-while-revalidate=30
# Strategy 3: explicit purge, only for what you cannot version.
# Shape varies by vendor, so check yours. Typical options:
# purge one URL | purge by prefix | purge by tag | purge everything
# Rule of thumb: purge the URL, then let the TTL clean up the rest.
Explicit purge is the most dangerous because it is manual and per PoP propagation is not instant, and the exact API differs by vendor, so read yours. Treat a purge as a way to speed up a short lifetime you already have, not as a substitute for setting one. A cache with long lifetimes and no versioning is one bad deploy away from serving stale pages you cannot fix quickly.
What is a quick checklist before putting a CDN in front of an app?
Run this list once per route group, not once per URL. It is short on purpose.
Split routes into public and personalised, and mark every personalised response private, no-store.
Put a content hash in the filename of every static asset and cache it for a year with immutable.
Give public HTML a short s-maxage with stale-while-revalidate, and max-age=0 for browsers.
Check the cache key: remove tracking parameters, avoid Vary: Cookie, and confirm what your vendor includes by default.
Read the Age header and your provider's hit or miss header on a real request before you trust any of it.
A CDN is HTTP caching with many caches in many places, so the rules are the ones in the RFCs: a clear key, explicit Cache-Control, versioned URLs, and a firm list of responses that never leave the origin. Decide what is public first, and the edge configuration mostly writes itself.