Short answers to what readers ask most about this topic.
01What is the difference between latency and throughput?
Latency is the time one request takes from start to finish, in milliseconds. Throughput is how many requests complete per unit of time, in requests per second. A system can have high throughput and poor latency at once, for example a batch pipeline that processes thousands of items per second but makes each wait a long time.
02What does Little's law say about latency and throughput?
Little's law states L = λW: the average number of requests in the system equals the arrival rate times the average time each request spends there. At 200 requests per second and 50 ms each, about 10 requests are in flight. It is useful for sizing connection pools and worker counts.
03Why does latency go up when a server is almost fully utilised?
Requests arrive at random, so some find the worker busy and must queue. In an M/M/1 queue the mean time in system is 1 / (μ − λ), a multiplier of 1 / (1 − ρ) over the service time. That is 5x at 80 percent utilisation, 10x at 90 percent and 100x at 99 percent.
04Why is p99 latency more important than average latency?
The average hides slow requests: one 2,400 ms request among ten fast ones pulls the mean to 252.4 ms, which describes none of them. Percentiles show how slow the worst experiences are. Tail latency also compounds when a request calls many backends, so 1 percent slow per backend becomes about 63 percent slow across 100.
05How do you reduce latency without lowering throughput?
Use techniques that help both: caching hot reads, connection pooling and keep-alive, removing unnecessary network hops, and moving non-essential work off the request path. Avoid fixes that trade one for the other, such as large batches or running near full utilisation. Always check the percentile before and after a change.
Latency vs Throughput: Differences, Trade-offs and Tuning
Latency vs throughput explained: what each measures, how Little's law links them, why batching and queues trade one for another, and how to optimise both.
Latency is how long one request takes; throughput is how many requests finish per unit of time. Little's law, L = λW, links them, yet they trade off: batching and high utilisation raise throughput while inflating latency, especially at p99. Cut latency by removing work and hops, and raise throughput with parallelism, pooling and capacity.
A dashboard says the API handles 800 requests per second, and a user says the checkout page feels slow. Both statements can be true at once, because they measure different things. Mixing them up is how teams spend a sprint adding servers to fix a delay that servers cannot fix.
This post defines both terms, connects them with Little's law, shows why they pull against each other using a batching table and an M/M/1 queue, and then explains percentiles and the optimisations that move each number. Every figure is either derived on the page or comes from my own small TypeScript simulation, which measures only that simulation and not any real server. For the network side of the same idea, see my earlier post on Jakarta vs Singapore PaaS latency, and for what to do when a queue fills, see the backpressure post.
What is the difference between latency and throughput?
Latency is the time one unit of work takes from start to finish, measured in milliseconds per request. Throughput is the amount of work completed per unit of time, measured in requests per second. One is a duration, the other is a rate, and the table shows how they differ in practice.
Aspect
Latency
Throughput
Question it answers
How long does one request take?
How many requests finish each second?
Unit
Milliseconds or seconds per request
Requests, transactions or bytes per second
Who feels it
The individual user waiting on a screen
The operator paying for capacity
How to report it
Percentiles: p50, p95, p99
Sustained average, plus the peak the system survives
Typical fix
Do less work per request, remove hops, cache
Parallelism, pooling, batching, more instances
A highway makes the point. Latency is how long one car takes to drive the road; throughput is how many cars pass a marker per hour. Adding lanes raises throughput and does nothing for a single car's trip time, while a lower speed limit hurts both.
How are latency and throughput connected by Little's law?
Little's law says the average number of requests in a stable system equals the arrival rate times the average time each spends there: L = λW. It holds regardless of the arrival distribution, which is why it is the first sanity check for any capacity question.
// Little's law: L = lambda * W (valid for any stable system, any distribution)
// L = average number of requests in the system (in flight)
// lambda = average arrival rate = throughput once stable (requests/second)
// W = average time each request spends in the system (seconds)
// 200 requests/s at 50 ms each:
// L = 200 * 0.050 = 10 requests in flight
// The same 200 requests/s after latency degrades to 250 ms:
// L = 200 * 0.250 = 50 requests in flight
// So a Postgres pool of max: 10 is enough in the first case and
// 40 requests are queueing for a connection in the second.
const pool = new Pool({ max: 10 });
The worked numbers show the consequence. At a steady 200 requests per second, a 50 ms latency keeps 10 requests in flight, but if latency drifts to 250 ms the same traffic keeps 50 in flight. Nothing about the traffic changed, yet the system now needs five times the concurrency to carry it, which is exactly how connection pools run dry.
Use L = λW backwards when sizing a pool: multiply your peak requests per second by the p95 latency in seconds, then add headroom. If the answer exceeds what Postgres or the downstream service can hold open, you have found the bottleneck on paper before production does.
Why does batching raise throughput but also latency?
Batching amortises a fixed per-call cost across many items, which is why bulk inserts and queue consumers that take many messages at once are so much faster in total. The catch is that an item must wait for its batch to fill. The table below uses a simple cost model I chose for illustration: 5 ms of fixed overhead per batch plus 0.5 ms per item, with one item arriving every 2 ms.
// Fixed 5 ms overhead per batch + 0.5 ms per item. One item arrives every 2 ms (500/s).
batch service capacity/s mean fill wait mean latency
1 5.5 ms 182 0.0 ms 5.5 ms
5 7.5 ms 667 4.0 ms 11.5 ms
10 10.0 ms 1000 9.0 ms 19.0 ms
20 15.0 ms 1333 19.0 ms 34.0 ms
50 30.0 ms 1667 49.0 ms 79.0 ms
Capacity climbs from 182 items per second at batch size 1 to 1,667 at batch size 50, a gain of about 9 times. Mean latency climbs from 5.5 ms to 79 ms over the same range, about 14 times, because the average item sits through half the fill time. Notice also that batch size 1 can only sustain 182 items per second against 500 arriving, so it would queue without bound; that is the case where batching is not optional.
The practical rule is to batch by size and by time together. Flush when the batch reaches a maximum size or when the oldest item has waited a maximum number of milliseconds, whichever comes first, so a quiet period does not strand items behind a batch that never fills.
Why does latency explode as utilisation rises?
For an M/M/1 queue, one worker with random arrivals and random service times, the mean time in system is W = 1 / (μ − λ). Dividing by the mean service time 1/μ gives the wait multiplier 1 / (1 − ρ), where ρ = λ/μ is utilisation. The multiplier is gentle until about 80 percent and then turns vertical, as the table shows for a 10 ms mean service time.
Utilisation ρ
Multiplier 1 / (1 − ρ)
Mean time W (derived)
Simulated p99
50%
2x
20 ms
93.6 ms
80%
5x
50 ms
230.7 ms
90%
10x
100 ms
438.5 ms
95%
20x
200 ms
838.2 ms
99%
100x
1,000 ms
4,497.9 ms
I checked the formula with a short simulation: Poisson arrivals, exponential service times, one worker, 400,000 jobs per level and a seeded random generator so the run repeats. The simulated mean matches the derived one within about 5 percent up to 95 percent utilisation, and at 99 percent it lands 7.6 percent high (1,075.5 ms against 1,000 ms) because a queue that close to saturation converges slowly.
// mm1.ts - run with: node mm1.ts (Node 22.18+ strips the types itself)
const SERVICE_MS = 10; // mean service time -> mu = 100 requests/s
const N = 400_000; // jobs per utilisation level
function simulate(rho: number) {
const rand = rng(42); // seeded PRNG, so the run is reproducible
const exp = (mean: number) => -mean * Math.log(1 - rand());
const meanInterarrival = SERVICE_MS / rho; // lambda = rho * mu
let wait = 0; // Lindley recursion: queueing delay of the current job
const sojourn: number[] = [];
for (let i = 0; i < N; i++) {
const service = exp(SERVICE_MS);
sojourn.push(wait + service); // time in system = queue wait + service
wait = Math.max(0, wait + service - exp(meanInterarrival));
}
sojourn.sort((a, b) => a - b);
return { mean: sojourn.reduce((s, x) => s + x, 0) / N, p99: percentile(sojourn, 99) };
}
// Output (ms), my own simulation of an idealised queue - not a measurement of any real server:
// rho theory W sim mean sim p50 sim p95 sim p99 p99/idle
// 0.50 20.0 20.1 13.9 60.4 93.6 9.4x
// 0.80 50.0 49.9 34.6 148.6 230.7 23.1x
// 0.90 100.0 97.6 68.9 283.1 438.5 43.8x
// 0.95 200.0 189.7 136.4 559.3 838.2 83.8x
// 0.99 1000.0 1075.5 684.1 3525.3 4497.9 449.8x
The simulation also shows why averages hide the pain. Waiting time in an M/M/1 queue is exponentially distributed, so p50 sits near 0.69 of the mean while p99 sits near 4.6 times the mean. At 90 percent utilisation the p99 is 438.5 ms for a job that takes 10 ms when the worker is idle. This is an idealised queue measured in a script, not a benchmark of Postgres, NestJS or any real server, but the shape of the curve is the part that transfers.
Running at 90 to 95 percent utilisation to save money is the expensive mistake. You pay for the spare capacity in latency instead of in rent, and a small bump in traffic moves you from 10x to 20x wait. Aim for sustained utilisation well below the knee and scale out before you reach it.
Why do p95 and p99 matter more than the average?
A percentile answers how slow the slowest requests are. The p50 is the value half of requests beat, p95 is the value 95 percent beat, and p99 is the value 99 percent beat. The snippet uses the nearest-rank method: sort the samples and take the value at rank ceil(p/100 x n). One slow request in ten drags the mean to 252.4 ms, a number that describes none of them.
// Nearest-rank percentile: sort ascending, take the value at rank ceil(p/100 * n).
function percentile(sorted: number[], p: number): number {
const rank = Math.ceil((p / 100) * sorted.length);
return sorted[Math.max(0, rank - 1)];
}
const latenciesMs = [12, 14, 13, 15, 12, 16, 13, 14, 15, 2400]; // one slow request in ten
const sorted = [...latenciesMs].sort((a, b) => a - b);
// Wrong: the mean says 252.4 ms, which describes none of the ten requests.
const mean = latenciesMs.reduce((s, x) => s + x, 0) / latenciesMs.length;
// Right: p50 = 14 ms is the typical request, p99 = 2400 ms is the slow one.
console.log(mean, percentile(sorted, 50), percentile(sorted, 99));
Tail latency also compounds. Google's Tail at Scale paper argues that at large fan-out the rare slow response becomes the common experience, and the arithmetic is short enough to do by hand. If each backend is slow 1 percent of the time, a request touching 100 of them is slow with probability 1 − 0.99^100, about 63.4 percent.
// A page that fans out to n backends is slow if ANY one of them is slow.
// If each backend is slow 1% of the time, P(page is slow) = 1 - 0.99^n
// n = 1 -> 1 - 0.99 = 1.0%
// n = 10 -> 1 - 0.9044 = 9.6%
// n = 100 -> 1 - 0.3660 = 63.4%
This is why a service-level objective should name a percentile and a threshold, for example 99 percent of checkouts under 800 ms, rather than an average. An average target can be met on a day when one in a hundred customers waits several seconds.
How do you optimise latency and throughput?
Start by deciding which number is the problem, because the fixes pull in different directions. Latency work removes time from a single request; throughput work adds capacity or removes per-item overhead. Measure the percentile first, then pick from the lists below.
To lower latency, do less per request:
Cache results close to the caller, such as Redis in front of a hot Postgres query, so most requests skip the slow path entirely.
Reduce hops: every extra service, proxy or region adds at least one round trip, so collapse chains and place services near each other.
Move non-essential work off the request path with async processing, so the user waits for the order to be saved but not for the email to be sent.
Reuse connections through keep-alive and pooling so each request avoids a fresh TCP and TLS handshake.
To raise throughput, increase the work done per unit of time:
Add parallelism: more worker processes or instances behind a load balancer, bounded by what the shared database can absorb.
Batch writes and queue reads so fixed costs are paid once per group instead of once per item.
Size connection pools from L = λW rather than guessing, since an undersized pool caps throughput and an oversized one overloads the database.
Keep utilisation below the knee so queueing delay does not eat the gain, and apply backpressure when it cannot be avoided.
Notice the overlap. Pooling and caching help both numbers, while batching and high utilisation buy throughput at the price of latency. Any change that improves one metric should be checked against the other before it ships.
What about bandwidth vs latency on the network?
Bandwidth is how many bits a link carries per second; latency is how long a bit takes to arrive. They are independent, and the bandwidth-delay product, bandwidth times round-trip time, tells you how much data must be in flight to fill a link. For small responses the round trip, not the bandwidth, sets the cost.
// Bandwidth-delay product = bandwidth x round-trip time = bytes in flight on the wire
// 100 Mbit/s link, 20 ms RTT:
// 100,000,000 bit/s * 0.020 s = 2,000,000 bits = 250,000 bytes (250 KB)
// A sender whose window is smaller than 250 KB cannot fill the link.
// A 2 KB JSON response, same link: transfer time = 16,000 bits / 100,000,000 = 0.16 ms.
// The 20 ms round trip dominates by roughly 125x. A faster link will not help;
// fewer round trips will (keep-alive, HTTP/2, a closer region).
The numbers above are plain arithmetic for a 100 Mbit/s link and a 20 ms round trip. They explain why moving a service closer to its users beats buying a faster uplink for chatty APIs, and why TCP window size matters on long, fast links. A 2 KB response spends 0.16 ms on the wire and 20 ms waiting for the round trip, so fewer round trips is the lever.
Treat latency and throughput as two dials on the same machine. Report latency in percentiles, size concurrency with L = λW, keep utilisation below the point where 1 / (1 − ρ) turns steep, and check every throughput optimisation for the latency it costs in return.