Short answers to what readers ask most about this topic.
01What is the bulkhead pattern and how does it contain failures?
The bulkhead pattern gives each dependency or workload its own fixed share of a resource, such as a connection pool, a set of concurrency slots, a queue or a container. When one dependency turns slow, it can only use up its own share. Everything else keeps its capacity, so the failure stays inside one compartment, like a flooded section of a ship.
02Why is it called the bulkhead pattern?
It is named after the sectioned partitions, called bulkheads, in a ship's hull. If the hull is breached, only the damaged section fills with water and the ship does not sink. Microsoft's Azure Architecture Center uses the same comparison and also calls the idea a cell-based architecture.
03How do I size a bulkhead?
Use Little's law: concurrency equals arrival rate times latency. A path that receives 40 requests per second and takes 0.05 seconds needs 40 x 0.05 = 2 slots. Then add headroom for bursts, and check that all your bulkheads together stay within the real limit of the database or service behind them.
04What is the difference between a bulkhead and a circuit breaker?
A bulkhead caps how much of your capacity one dependency can consume while it is slow, using a static limit on concurrent calls. A circuit breaker stops calling a dependency after it has been observed failing, and tests it again later. They solve different problems, and the Azure pattern guidance recommends combining them with retry and throttling.
05Should a full bulkhead queue requests or reject them?
Reject them, or allow only a very short queue. Queueing turns overload into latency, which keeps the caller's own threads and connections held open and spreads the problem. Returning something like HTTP 503 immediately lets the client back off or show a clear message.
Bulkhead Pattern Explained: Isolate Failures in Backend Systems
What the bulkhead pattern is, how one slow dependency drains a shared pool, how to isolate with pools, semaphores, queues and containers, and how to size each one.
The bulkhead pattern partitions resources such as connection pools, concurrency slots, queues or containers per dependency, so one slow or failing dependency can exhaust only its own partition. Size each partition with Little's law, concurrency equals arrival rate times latency, and reject excess calls immediately instead of queueing them.
The outage I worry about most in a point-of-sale or carwash ERP is not a crash. A crash is loud and gets fixed. It is the quiet one: a slow monthly report query holds database connections, the checkout endpoint waits behind it for a connection of its own, and the cashier sees a spinner on a screen that has nothing to do with reports.
This post explains the bulkhead pattern as the answer to that failure. I have not benchmarked anything for it. The pattern and its options come from the Azure Architecture Center, the pool behaviour from the node-postgres documentation, the container flags from the Docker documentation, and the sizing from Little's law. The only measured output is the counts printed by a TypeScript demo I ran, and the other numbers are arithmetic on stated assumptions.
What is the bulkhead pattern?
A bulkhead isolates the parts of an application into separate pools, so that if one part fails the others keep working. The name comes from the sectioned partitions in a ship's hull: if the hull is breached, only the damaged section fills with water, and the ship stays afloat. Microsoft's Azure Architecture Center describes it this way and notes that the same idea is also called a cell-based architecture.
The unit being protected is a resource, not a request. A thread, a database connection, an in-flight HTTP call or a slice of memory is finite, and a bulkhead gives each dependency its own fixed share of it. A failure then stays inside the share it started in. This is a different job from retrying or timing out: those decide what one call does, while a bulkhead decides how much of the system one dependency is allowed to take.
How does one slow dependency exhaust a shared pool?
Take a NestJS API on one VPS where every query goes through one node-postgres Pool. The documentation gives max a default of 10 clients, and says that when the pool is full and all clients are checked out, requests wait in a FIFO queue until a client is released. That queue is the problem: a slow caller does not fail, it makes everyone else join the line.
Little's law makes the damage countable. It states L = λW, where L is the average number of items in the system, λ is the arrival rate and W is the average time each spends in it. Assume checkout runs at 40 requests per second and holds a connection for 0.05 seconds, so it needs 40 × 0.05 = 2 connections. Assume reports arrive at 2 per second and take 1.5 seconds, so they need 2 × 1.5 = 3. Together that is 5 of 10 connections, a comfortable pool.
Now the report query degrades and takes 30 seconds. Reports now want 2 × 30 = 60 connections, but the pool only has 10. Arrivals of 2 per second fill the 10 slots in about 5 seconds, and from then on checkout waits in the same queue behind reports. Nothing crashed and no error was thrown. Checkout is down because it shares a pool, which is exactly what a bulkhead prevents.
How do you implement a bulkhead in TypeScript?
The smallest useful bulkhead is a semaphore with an admission rule: count the calls in flight, and when the count reaches the limit, either let a bounded number wait or throw at once. The class below does that. When a task finishes it hands its slot directly to the next waiter, so the active count never drops to zero and re-rises between them. With the default maxQueue of 0 it never waits at all.
export class BulkheadFullError extends Error {
readonly bulkhead: string;
constructor(bulkhead: string) {
super(`bulkhead "${bulkhead}" is full`);
this.bulkhead = bulkhead;
}
}
export class Bulkhead {
private active = 0;
private readonly waiters: Array<() => void> = [];
private readonly name: string;
private readonly maxConcurrent: number;
private readonly maxQueue: number;
// maxQueue 0 means: reject the moment every slot is taken.
constructor(name: string, maxConcurrent: number, maxQueue = 0) {
this.name = name;
this.maxConcurrent = maxConcurrent;
this.maxQueue = maxQueue;
}
async run<T>(task: () => Promise<T>): Promise<T> {
if (this.active < this.maxConcurrent) {
this.active++;
} else if (this.waiters.length < this.maxQueue) {
// The releasing task hands its slot straight to us, so `active` stays the same.
await new Promise<void>((resolve) => this.waiters.push(resolve));
} else {
throw new BulkheadFullError(this.name); // fail fast: nothing waits, no thread is held
}
try {
return await task();
} finally {
const next = this.waiters.shift();
if (next) next(); // pass the slot on
else this.active--; // or free it
}
}
}
To see the isolation, the demo gives the same 7 slots two layouts. In the first, reports and payments share one bulkhead of 7. In the second, reports get 3 and payments get 4. Each scenario fires 10 report calls that hang on a gate, then 4 payment calls that answer immediately. Only counts are printed, because the point is who gets admitted, not how fast.
const TOTAL_SLOTS = 7;
const REPORT_CALLS = 10; // slow dependency: every call hangs until the gate opens
const PAYMENT_CALLS = 4; // fast dependency: answers immediately
type Tally = { accepted: number; rejected: number };
function fire(box: Bulkhead, task: () => Promise<unknown>, n: number, tally: Tally) {
return Array.from({ length: n }, () =>
box.run(task).then(
() => void tally.accepted++,
(err) => {
if (!(err instanceof BulkheadFullError)) throw err;
tally.rejected++;
},
),
);
}
async function scenario(label: string, reports: Bulkhead, payments: Bulkhead) {
let open!: () => void;
const gate = new Promise<void>((resolve) => (open = resolve));
const reportTally: Tally = { accepted: 0, rejected: 0 };
const paymentTally: Tally = { accepted: 0, rejected: 0 };
const calls = [
...fire(reports, () => gate, REPORT_CALLS, reportTally), // hang, holding their slots
...fire(payments, async () => "ok", PAYMENT_CALLS, paymentTally),
];
open(); // let the hung report calls finish so the script can exit
await Promise.all(calls);
console.log(label);
console.log(` reports : sent ${REPORT_CALLS}, accepted ${reportTally.accepted}, rejected ${reportTally.rejected}`);
console.log(` payments: sent ${PAYMENT_CALLS}, accepted ${paymentTally.accepted}, rejected ${paymentTally.rejected}`);
}
const shared = new Bulkhead("shared", TOTAL_SLOTS);
await scenario(`One shared bulkhead of ${TOTAL_SLOTS}`, shared, shared);
await scenario(
"Two bulkheads: reports 3, payments 4",
new Bulkhead("reports", 3),
new Bulkhead("payments", 4),
);
The output is real, from running the file with Node 26.10.0. With one shared bulkhead the 10 hung report calls took all 7 slots and rejected 3, which left no slot for payments: all 4 were rejected. With the split, reports were capped at 3 and rejected 7 of their own calls, while all 4 payments were accepted. The total capacity was identical, only its partitioning changed.
$ node bulkhead.ts
One shared bulkhead of 7
reports : sent 10, accepted 7, rejected 3
payments: sent 4, accepted 0, rejected 4
Two bulkheads: reports 3, payments 4
reports : sent 10, accepted 3, rejected 7
payments: sent 4, accepted 4, rejected 0
Which resources can you partition into bulkheads?
The Azure pattern page lists several levels of isolation, and they trade strength against cost. The table orders them from cheapest to strongest. Pick the cheapest level that the failure you fear cannot cross.
Level
What it isolates
Cost
Reach for it when
Separate connection pools
Database or HTTP connections per dependency
Almost free, but the pool totals must stay under the server limit
One database serves workloads of very different weight
Concurrency semaphore
In-flight calls inside one process
A few lines of code, no process overhead
You call several downstream services from one process
Separate queues and workers
Asynchronous work such as emails, exports, webhooks
A queue per class of work, plus its workers
A backlog in one job type must not delay another
Separate processes or containers
CPU, memory and crashes
More memory overhead, more things to deploy and monitor
A runaway report can starve the whole machine
On a single VPS the container level is where the hardware limits get enforced. Docker documents memory as the maximum a container can use, and --cpus as how much of the available CPU a container may use, equivalent to setting the CPU period and quota together. Setting --memory-swap equal to --memory disables swap for the container, which stops a leaking report worker from quietly pushing the host into swap. The sizes below are illustrative.
# Separate containers, each with its own ceiling. Sizes are illustrative.
# --memory-swap equal to --memory disables swap for the container (Docker docs).
# --cpus=0.5 caps the container at half of one CPU.
docker run -d --name api --memory=512m --memory-swap=512m --cpus=1.0 myorg/api
docker run -d --name report-worker --memory=256m --memory-swap=256m --cpus=0.5 myorg/report-worker
The Azure page also gives a Kubernetes example with requests and limits for memory and CPU, and observes that containers offer a good balance of isolation and low overhead. It adds a warning worth taking seriously: where a platform already provides throttling and isolation, such as API gateway rate limits, use those instead of rebuilding them in application code.
How big should each bulkhead be?
Start from Little's law and work backwards. The steady-state need of a dependency is its arrival rate times its latency, which gave 2 connections for checkout and 3 for reports above. A bulkhead set exactly at that number rejects calls at every small burst, so add headroom. The factor is a judgement, not a law: I would give a fast, critical path about three times its need and a slow, optional one only a little above it.
Applied to the earlier 10-connection budget, checkout gets 6 (three times its need of 2) and reports get 4 (a third above its need of 3), and 6 + 4 = 10, so the pools together stay inside the budget. If reports now degrade to 30 seconds, at most 4 report calls are in flight and the rest are rejected, while checkout keeps its own 6. Reports are worse off and checkout is untouched, which is the intended trade.
The two pools in node-postgres are separate Pool objects, each with its own max. Read the pool waitingCount property, which the documentation exposes for the number of queued requests, because a pool that regularly has waiters is sized too small for its traffic.
import { Pool } from "pg";
// Wrong: one pool, default max of 10, shared by checkout and reports.
// export const db = new Pool();
// Right: two pools whose max values add up to the budget of 10.
export const checkoutPool = new Pool({ max: 6 }); // need 2 (40/s x 0.05 s), headroom x3
export const reportPool = new Pool({
max: 4, // need 3 (2/s x 1.5 s), a third above it
connectionTimeoutMillis: 2000, // docs: the default 0 means no timeout
});
// pool.waitingCount is the number of queued requests: export it as a metric.
export const poolStats = () => ({
checkoutWaiting: checkoutPool.waitingCount,
reportWaiting: reportPool.waitingCount,
});
Size the sum of all bulkheads against the real limit behind them, not each in isolation. If four pools of 10 point at a database that accepts fewer than 40 connections, the bulkheads are decoration and the database becomes the shared pool.
Should a full bulkhead queue or reject?
Reject, or queue only a very short line. A queue converts overload into latency, and latency is what holds the caller's own resources open. The Resilience4j documentation reflects the same bias: its semaphore bulkhead defaults to maxConcurrentCalls of 25 and a maxWaitDuration of 0, so a saturated bulkhead refuses immediately unless you configure a wait.
In the API layer a rejection should become an honest, cheap response. HTTP 503 means the service cannot handle the request right now, and the caller or the user interface can then show a clear message or try again shortly, instead of freezing. The code below wraps a report query in its own bulkhead and maps the full condition to a 503.
import { Inject, Injectable, ServiceUnavailableException } from "@nestjs/common";
import type { Pool } from "pg";
import { Bulkhead, BulkheadFullError } from "./bulkhead";
@Injectable()
export class ReportService {
// 4 slots, no queue: the fifth concurrent report is refused, not parked.
private readonly box = new Bulkhead("reports", 4);
constructor(@Inject("REPORT_POOL") private readonly pool: Pool) {}
async monthly(branchId: number) {
try {
return await this.box.run(() =>
this.pool.query("SELECT * FROM monthly_sales($1)", [branchId]),
);
} catch (err) {
if (err instanceof BulkheadFullError) {
// 503: the service cannot take this right now. Cheap, immediate, honest.
throw new ServiceUnavailableException("Report capacity reached, retry shortly");
}
throw err;
}
}
}
Do not retry a rejected call in a tight loop. The rejection says the partition is full, and instant retries from every client rebuild the pressure the bulkhead just removed. Retry with backoff and jitter, or let the user press the button again.
How do you observe a bulkhead?
A bulkhead that nobody watches fails silently, because its rejections look like ordinary errors. Track four numbers per bulkhead: calls in flight against the limit, calls waiting, rejections per interval, and the latency of admitted calls. In the demo class that means exposing the active count and a rejected counter, and for node-postgres it means the pool total and waitingCount.
Alert on the rejection rate and on sustained saturation, not on a single spike. Rising rejections with normal admitted latency means the limit is too low or demand grew. Rising rejections with rising admitted latency means the dependency itself is degrading, and that is the signal to look at the dependency, not at the bulkhead. The Azure page makes the same recommendation to monitor each partition's performance and service-level agreement.
Bulkhead vs circuit breaker: when do you use which?
They protect against different things and are meant to be combined, which the Azure page states directly. A bulkhead limits how much a dependency can consume while it is slow. A circuit breaker stops calling a dependency after it has proved to be failing. My post on the circuit breaker pattern in NestJS covers the second; the table compares the two.
A static limit on concurrent calls or a queue length
An observed failure or timeout rate over a window
Protects against
Slowness that holds resources, and noisy neighbours
Repeated calls to something that is already failing
Stateful
Only a count of calls in flight
Yes: closed, open and half-open states
Use a bulkhead whenever two workloads of different importance share a limited resource. Use this checklist to decide where the partitions go.
List the finite resources behind the service: database connections, outbound HTTP calls, CPU and memory, worker slots.
Mark each workload as critical (checkout, payments, login) or optional (reports, exports, analytics).
Compute each workload's need as arrival rate times latency, then add headroom.
Make sure the sum of the limits does not exceed the real limit of the resource behind them.
Set rejection rather than waiting as the default, return 503, and alert on the rejection rate.
A shared pool is a promise that nobody will hog it, and the slowest dependency is the one that breaks that promise. Give each dependency its own share, size the share with arrival rate times latency, fail fast when it fills, and watch the rejection rate. The failure then stays in the compartment where it began.