Short answers to what readers ask most about this topic.
01How should you retry failed requests?
Retry only transient failures such as timeouts, connection resets, 429 and 503, and only for idempotent operations or requests that carry an idempotency key. Wait with capped exponential backoff plus jitter, honour any Retry-After header, and limit total attempts with a budget and a deadline. Do not retry other 4xx errors, because the same request fails the same way.
02What is exponential backoff with jitter?
Exponential backoff doubles the wait after each failed attempt up to a cap. Jitter adds randomness to that wait so many clients do not retry at the same instant. With full jitter, each client sleeps a random time between zero and the capped exponential value.
03Why is jitter needed if you already use exponential backoff?
Without jitter, clients that failed together wait the same doubling delays and retry together, so the load spike repeats every round. In a simulation of 1000 clients failing at the same moment with a 50 ms window, no jitter put all 1000 retries in one window every round, while full jitter spread the first round across 4 windows and the fourth across 52. The AWS Architecture Blog also calls backoff without jitter the clear loser.
04Which HTTP errors are safe to retry?
Timeouts, connection resets, 408, 429, 502, 503 and 504 are transient and worth retrying, provided the operation is idempotent. For 429 and 503, wait at least as long as Retry-After says when it is present. Errors such as 400, 401, 403, 404 and 422 mean the request is wrong, so repeating it only adds load.
05How do you stop retries from causing a retry storm?
Retry at one layer only, since three layers each making 3 tries multiply to 27 calls at the bottom. Add a retry budget, such as the Google SRE book's limit of three attempts per request and retries below 10% of requests per client. Back off with jitter and set an overall deadline so retries stop when they can no longer help.
Retry with Exponential Backoff and Jitter in TypeScript
How to retry failed requests safely: which errors to retry, the retry storm arithmetic, exponential backoff with jitter, Retry-After, retry budgets and idempotency keys, with a runnable TypeScript helper.
Retry only transient failures on idempotent operations: timeouts, connection resets, 429 and 503, never other 4xx errors. Wait with capped exponential backoff plus full jitter so clients spread out instead of retrying in lockstep. Honour Retry-After, retry at one layer only, cap attempts with a budget and a deadline, and send an idempotency key on writes.
Picture a fleet of POS terminals and a payment provider that returns 503 for ten seconds. Every terminal has a retry loop, so every terminal retries, and a short outage becomes a wall of traffic the provider has to survive just as it comes back. A retry is a bet that the next attempt will succeed. Placed carelessly, it is also extra load aimed at the thing that is already failing.
This post answers one question: how should you retry failed requests? It covers which errors deserve a retry, how retries multiply across layers, how exponential backoff and jitter spread them out, and what Retry-After, budgets, deadlines and idempotency keys add. The rules come from the AWS Architecture Blog, the Google SRE book and two RFCs. The retry helper and the simulation are code I ran, and the output is pasted unedited. The circuit breaker and idempotency key posts on this site cover two neighbouring topics in depth.
Which errors should you retry, and which should you never retry?
Retry failures that are transient and operations that are safe to repeat. Both conditions matter. A timeout is transient, but if the request was a payment, you do not know whether the server already did the work. RFC 9110 defines a method as idempotent when repeating it has the same intended effect as sending it once, and lists PUT, DELETE and the safe methods as idempotent. POST is not on that list.
Signal
Retry?
Why
Timeout or connection reset
Yes, if idempotent or sent with an idempotency key
You cannot tell whether the server processed the request before the connection died.
429 Too Many Requests
Yes, after the wait the server asks for
RFC 6585 defines it as rate limiting and allows a Retry-After header saying how long to wait.
503 Service Unavailable
Yes, honouring Retry-After
RFC 9110 says Retry-After on a 503 indicates how long the service is expected to be unavailable.
502 or 504 from a proxy
Yes, if idempotent
The proxy failed, and the origin behind it may or may not have handled the request.
400, 401, 403, 404, 422
No
The request itself is wrong. Sending the same bytes again fails the same way and only adds load.
A plain 500 sits in between. It is often a bug in the handler, which a retry will hit again, so retry it at most once and only when the operation is idempotent. The 401 case is the one exception to the 4xx rule: refresh the token, then send a new request. That is a different action from retrying, so it should not consume an attempt from the retry budget.
What is a retry storm, and how do retries multiply across layers?
A retry storm is load created by the retries themselves. The cause is layering: a gateway calls orders, orders calls payments, payments calls a bank API. If each layer makes up to 3 tries, a call that keeps failing is tried 3 times at each layer, and the multiplication happens below the first layer. The Google SRE book names this a combinatorial explosion. The arithmetic is below, and it is derived, not measured.
Gateway -> Orders -> Payments -> Bank API
3 tries 3 tries 3 tries (every layer retries a failed call up to 2 times)
Calls reaching the Bank API per user request when it keeps failing:
1 layer retrying: 3
2 layers retrying: 3 x 3 = 9
3 layers retrying: 3 x 3 x 3 = 27
4 layers retrying: 3 x 3 x 3 x 3 = 81
A budget of 10% retries per client instead (SRE book): about 1.1x, not 27x.
The SRE book's remedy has two parts. First, retry only at the layer immediately above the one rejecting the request, so the layers do not multiply. Second, put a budget on retries. It describes a per-request limit of up to three attempts and a per-client limit that allows a retry only while retries are below 10% of requests. Without the client limit it describes traffic growing to just below 3x the original rate, and with it, about 1.1x. When most tasks are overloaded, the book says to stop retrying and let the error bubble up to the caller.
Layered retries hide in libraries. An HTTP client, a message queue consumer and an API gateway can each retry by default, and nobody sees the product until the downstream service falls over. List every layer between the user and the failing dependency, and decide which single one retries.
How does exponential backoff work, and why is it not enough?
Exponential backoff doubles the wait after each failed attempt, up to a cap. With a base of 100 ms, the delays are 200, 400, 800 and 1600 ms for the first four retries in the helper below, and the cap keeps a long outage from producing hour-long sleeps. The idea is old, and the Wikipedia article on exponential backoff traces it to collision handling in shared networks. The point is that a client which keeps failing should ask less and less often.
The flaw is that every client doubles in step. If a thousand clients fail at the same moment, they all wait exactly 200 ms, all retry together, fail together, and all wait exactly 400 ms. The AWS Architecture Blog post by Marc Brooker calls plain exponential backoff without jitter the clear loser, needing more work and more time than the jittered variants. The delays grow, but the spike does not spread out, so the server sees one wall of requests per round.
Which jitter should you use: full, equal or decorrelated?
Jitter adds randomness to the wait so clients stop retrying in lockstep. The AWS post compares three versions. Full jitter picks a random delay between zero and the capped exponential value. Equal jitter keeps half of the backoff and randomises the other half. Decorrelated jitter bases the next maximum on the previous sleep instead of the attempt count.
Strategy
Sleep before retry n
What stands out
No jitter
min(cap, base x 2 to the n)
Every client wakes at the same instant. AWS calls it the clear loser.
Full jitter
random from 0 to min(cap, base x 2 to the n)
Widest spread of the capped variants, but a client can draw a sleep close to zero.
Equal jitter
half the backoff plus random up to the other half
Guarantees a minimum wait. AWS reports slightly more work than full jitter and much longer completion.
Decorrelated jitter
min(cap, random from base to previous sleep x 3)
Grows from the last sleep. AWS reports it makes more calls than full jitter.
To see the clustering rather than assert it, I simulated 1000 clients that all fail at t=0 and counted how many of their retries land in each 50 ms window. The run uses the same backoffDelay function as the helper in the next section, a base of 100 ms, a cap of 10 seconds and a seeded random generator so it repeats. It measures clustering only. It says nothing about how any real server would perform.
// simulate.ts: 1000 clients all fail at t=0. For retries 1 to 4, each client sleeps
// backoffDelay(...) after its own previous retry; count how many land in each 50 ms window.
// base = 100 ms, cap = 10000 ms, seeded PRNG (mulberry32, seed 42) so the run repeats.
for (const strategy of ["none", "full", "equal", "decorrelated"]) { /* ... */ }
$ node simulate.ts
clients=1000 all fail at t=0, base=100ms cap=10000ms, window=50ms, seed=42
none
retry 1: occupied windows=1, peak retries in one window=1000
retry 2: occupied windows=1, peak retries in one window=1000
retry 3: occupied windows=1, peak retries in one window=1000
retry 4: occupied windows=1, peak retries in one window=1000
full
retry 1: occupied windows=4, peak retries in one window=277
retry 2: occupied windows=12, peak retries in one window=147
retry 3: occupied windows=26, peak retries in one window=72
retry 4: occupied windows=52, peak retries in one window=48
equal
retry 1: occupied windows=2, peak retries in one window=520
retry 2: occupied windows=6, peak retries in one window=265
retry 3: occupied windows=14, peak retries in one window=126
retry 4: occupied windows=26, peak retries in one window=79
decorrelated
retry 1: occupied windows=4, peak retries in one window=277
retry 2: occupied windows=20, peak retries in one window=124
retry 3: occupied windows=57, peak retries in one window=52
retry 4: occupied windows=116, peak retries in one window=33
Without jitter, all 1000 retries land in a single window in every round. With full jitter, the first retry round spreads across 4 windows with a peak of 277, and by the fourth round across 52 windows with a peak of 48. Equal jitter is tighter, at 2 windows and a peak of 520 in the first round. Decorrelated jitter ends up spread widest by round four, across 116 windows. The AWS post adds a caveat worth keeping: jitter reduces the work, but none of these variants change the quadratic nature of many clients contending for one server.
Start with full jitter. It is one line, it needs no per-client state, and AWS reports it making fewer calls than decorrelated jitter. Switch only if you have measured a reason, and keep the cap, because jitter does not bound the wait.
How do Retry-After, deadlines and retry budgets fit into one helper?
The helper below combines the pieces. It retries only the errors in the table, picks a strategy, treats Retry-After as a floor, gives each attempt only what remains of an overall deadline, and gives up rather than sleep past that deadline. RFC 9110 says Retry-After is either a number of seconds or an HTTP-date, so parseRetryAfter handles both. It runs as plain TypeScript on Node 22.18 or later, where the types are stripped without a build step.
export type Strategy = "none" | "full" | "equal" | "decorrelated";
export interface RetryOptions {
maxAttempts: number; // total tries, including the first
baseMs: number;
capMs: number;
deadlineMs: number; // wall-clock budget for the whole call, retries included
strategy: Strategy;
rng?: () => number;
sleep?: (ms: number) => Promise<void>;
}
export class HttpError extends Error {
status: number;
retryAfterMs?: number;
constructor(status: number, retryAfterMs?: number) {
super("HTTP " + status);
this.status = status;
this.retryAfterMs = retryAfterMs;
}
}
// Transient by nature. 400, 401, 403, 404, 422 are absent on purpose: repeating them fails again.
const RETRYABLE_STATUS = new Set([408, 429, 502, 503, 504]);
export function isRetryable(err: unknown): boolean {
if (err instanceof HttpError) return RETRYABLE_STATUS.has(err.status);
const code = (err as { code?: string }).code;
return code === "ECONNRESET" || code === "ETIMEDOUT" || code === "ECONNREFUSED";
}
// Retry-After is either delay-seconds or an HTTP-date (RFC 9110 section 10.2.3).
export function parseRetryAfter(value: string | null, now = Date.now()): number | undefined {
if (!value) return undefined;
if (/^\d+$/.test(value)) return Number(value) * 1000;
const at = Date.parse(value);
return Number.isNaN(at) ? undefined : Math.max(0, at - now);
}
// attempt is 1 for the first retry. prev is the previous sleep (decorrelated only).
export function backoffDelay(
o: Pick<RetryOptions, "baseMs" | "capMs" | "strategy">,
attempt: number,
prev: number,
rng: () => number,
): number {
const exp = Math.min(o.capMs, o.baseMs * 2 ** attempt);
switch (o.strategy) {
case "none":
return exp;
case "full":
return rng() * exp;
case "equal":
return exp / 2 + (rng() * exp) / 2;
case "decorrelated":
return Math.min(o.capMs, o.baseMs + rng() * (prev * 3 - o.baseMs));
}
}
export async function retry<T>(
op: (attempt: number, signal: AbortSignal) => Promise<T>,
o: RetryOptions,
): Promise<T> {
const rng = o.rng ?? Math.random;
const sleep = o.sleep ?? ((ms: number) => new Promise<void>((r) => setTimeout(r, ms)));
const started = Date.now();
let prev = o.baseMs;
for (let attempt = 1; ; attempt++) {
const left = o.deadlineMs - (Date.now() - started);
if (left <= 0) throw new Error("deadline exceeded before attempt " + attempt);
try {
// Each attempt gets only what is left of the overall deadline.
return await op(attempt, AbortSignal.timeout(left));
} catch (err) {
if (!isRetryable(err) || attempt >= o.maxAttempts) throw err;
let wait = backoffDelay(o, attempt, prev, rng);
prev = wait;
// The server's hint is a floor: never retry sooner than it asked.
const hint = err instanceof HttpError ? err.retryAfterMs : undefined;
if (hint !== undefined) wait = Math.max(wait, hint);
if (wait >= o.deadlineMs - (Date.now() - started)) throw err; // would wake after the deadline
await sleep(wait);
}
}
}
// Run with: node smoke.ts (Node 22.18+ strips the types; no build step)
const slept: number[] = [];
let calls = 0;
const result = await retry(
async () => {
calls++;
if (calls < 3) throw new HttpError(503, calls === 1 ? 1500 : undefined);
return "ok";
},
{ maxAttempts: 5, baseMs: 100, capMs: 10_000, deadlineMs: 30_000,
strategy: "full", rng: () => 0.5, sleep: async (ms) => { slept.push(ms); } },
);
console.log(result, "calls=" + calls, "sleeps=" + JSON.stringify(slept));
// Real output:
// ok calls=3 sleeps=[1500,200]
// retry 1: rng 0.5 x 200 ms = 100 ms, but Retry-After said 1500 ms, so it waited 1500
// retry 2: rng 0.5 x 400 ms = 200 ms, no hint
// A 404 is not retried at all:
// 404 calls=1 Error: HTTP 404
I ran a smoke test with a fixed random value of 0.5 and a fake sleep. A 503 carrying a 1500 ms hint followed by a plain 503 and then a success took three calls and slept 1500 then 200 ms: the first computed delay was 100 ms, but the server's hint wins. A 404 made exactly one call and was thrown at once. The helper has no retry budget across calls. A shared counter of retries against requests, as in the SRE book, belongs in the client that owns the connection pool, not in each call site.
How do idempotency keys make retries safe?
An idempotency key turns a non-idempotent write into one that is safe to repeat. The client generates one key per logical operation, sends it with every attempt, and the server returns the stored result for a key it has already seen instead of doing the work twice. The detail that matters is where the key is created. It has to exist before the first attempt and be reused by every retry.
// One key per LOGICAL operation, created before the first attempt and reused on every retry.
const key = crypto.randomUUID();
await retry(
(attempt, signal) =>
fetchOrThrow("https://api.example.com/v1/payments", {
method: "POST", // not idempotent by itself; the key makes the retry safe
headers: { "Idempotency-Key": key, "Content-Type": "application/json" },
body: JSON.stringify({ orderId: "ord_1042", amountIdr: 85000 }),
signal,
}),
{ maxAttempts: 3, baseMs: 200, capMs: 5_000, deadlineMs: 10_000, strategy: "full" },
);
// Wrong: a new key inside the callback. Every retry is then a brand-new payment.
If the key is generated inside the retried callback, each attempt is a new operation and a timeout followed by a retry can charge the customer twice. The key also stops the dangerous ambiguity from the first table: after a timeout you can retry a payment without knowing whether it went through. The idempotency keys post on this site covers the server side, storing the key and its response, and the circuit breaker post covers what to do when retries keep failing and you should stop calling at all.
What is the checklist before you ship a retry?
Run through these in order. Each one removes a failure mode from the sections above, and skipping any of them is how a helpful retry becomes an outage amplifier.
Retry only timeouts, connection resets, 408, 429, 502, 503 and 504, and never the other 4xx errors.
Retry only idempotent operations, or attach an idempotency key created before the first attempt.
Use capped exponential backoff with full jitter, and treat Retry-After as a minimum wait.
Pick one layer to retry and turn retries off in every other layer between the user and the dependency.
Cap attempts, and bound retries with a budget such as 10% of requests, so a long outage cannot multiply traffic.
Set an overall deadline, pass the remaining time to each attempt, and stop when the next sleep would pass it.
The rule I carry away is that a retry is a request for more capacity from a service that is already struggling, so it has to be rare, spread out and safe to repeat. Retry the right errors, add jitter, respect Retry-After, retry in one place, set a budget and a deadline, and make writes idempotent. Done that way, a brief outage stays brief instead of becoming a storm.