Short answers to what readers ask most about this topic.
01Which load balancing algorithm is best?
No single algorithm wins; it depends on what varies in your traffic. Round robin suits identical servers with uniform short requests, least connections suits mixed request durations or long-lived connections, and hashing suits cases where a client must return to the same server. Start with round robin or least_conn and change only when you can show the simple choice misbehaves.
02What is the difference between round robin and least connections?
Round robin rotates through servers and counts requests, so it is blind to how long each request takes. Least connections sends the next request to the server with the fewest active connections, which tracks real load more closely when request durations differ. In nginx the second is enabled with the least_conn directive inside the upstream block.
03Does ip_hash give you sticky sessions?
Mostly. The nginx docs say ip_hash sends requests from the same client to the same server unless that server is unavailable, using the first three octets of an IPv4 address as the key. That means a whole /24 shares one server and clients behind a proxy or NAT all look alike. Storing sessions in Redis avoids the need for stickiness.
04What is the difference between layer 4 and layer 7 load balancing?
Layer 7 balancing happens after the balancer has parsed the HTTP request, so it can use URLs, headers and cookies and terminate TLS. Layer 4 balancing works on the TCP connection and cannot see the request. In nginx the first is the http block and the second is the stream block, which offers hash, least_conn, least_time and random.
05What happens to load balancing when a backend server goes down?
That depends on health checks. Open-source nginx uses passive checks: after max_fails failed responses (default 1) it avoids the server for fail_timeout (default 10 seconds), then probes it with live requests. Until the failure is detected, round robin and ip_hash keep sending traffic to the dead server, while least connections tends to drift away from a slow one.
Use round robin for identical servers handling short, similar requests, least connections when request durations vary, and a hash of a client key (ip_hash or hash with consistent) only when one client must reach the same server. Add weights for unequal machines, and pair every algorithm with health checks so failed servers stop receiving traffic.
Adding a second replica of an API container behind nginx forces a decision that the default config makes silently: which server gets the next request. Nginx answers with round robin unless told otherwise, and for a while that answer is good enough that nobody checks it.
This post is about the algorithms themselves and when each one fits: round robin, least connections, IP hash and consistent hashing, weights, and power of two choices. Each gets an nginx upstream snippet, the layer 4 versus layer 7 difference is covered, and so are health checks. Every behaviour claim traces to the nginx documentation, the HAProxy engineering blog or Wikipedia, and every number is either quoted from them or worked out on the page. For the surrounding config, see the separate post on a production nginx load balancer.
What does a load balancing algorithm actually decide?
A load balancer does one job per request or per connection: pick one server from the pool. The algorithm is only the rule for that pick. Wikipedia groups the common rules as round robin, weighted round robin, least connections and hash-based assignment, and notes that the first two can be weighted so stronger machines receive more.
The rules differ in what they need to know. Round robin needs nothing but a counter. Least connections needs the live number of open connections per server. A hash needs a key from the request, such as the client address or a cookie. The more a rule knows, the better it adapts, and the more state the balancer has to keep.
When is round robin enough?
Round robin sends request 1 to server 1, request 2 to server 2, and wraps after the last server. The nginx load balancing guide states that when no method is configured it defaults to round robin, and it attaches the condition that matters: the split is even only when requests are processed uniformly and finish fast enough.
# Round robin is the default: no method directive at all.
upstream api_pool {
server 10.0.0.11:3000;
server 10.0.0.12:3000;
server 10.0.0.13:3000;
}
server {
listen 80;
location / {
proxy_pass http://api_pool;
}
}
That makes it the right default for identical, stateless replicas serving requests of similar cost, such as a NestJS API where nearly every endpoint is a short indexed Postgres query. It stops being right as soon as request cost varies a lot, which the next section shows with arithmetic.
When should you use least connections instead?
Least connections sends the next request to the server with the fewest active connections, and nginx takes server weights into account, falling back to weighted round robin when several servers tie. The nginx guide recommends it when some requests take longer, so a busy server is not given excessive extra work.
Here is a worked example of why. Take two servers and eight requests that alternate heavy (900 ms) and light (100 ms), starting with a heavy one. Strict alternation puts every heavy request on the same server.
# 8 requests, alternating heavy (900 ms) and light (100 ms)
# Round robin over servers A and B:
# A gets requests 1,3,5,7 (all heavy): 4 x 900 = 3600 ms of work
# B gets requests 2,4,6,8 (all light): 4 x 100 = 400 ms of work
# total 4000 ms, fair share 2000 ms each, A carries 90 percent
# Conceptually, least_conn picks the server with the fewest open
# connections (weights considered, ties broken by weighted round robin).
# A, stuck on a 900 ms request, keeps its connection open, so the next
# request goes to B, which finishes its 100 ms ones quickly.
upstream api_pool {
least_conn;
server 10.0.0.11:3000;
server 10.0.0.12:3000;
}
The pattern is deliberately extreme, but the mechanism is real: round robin counts requests, not work. Least connections measures something closer to work, because a server stuck on a slow request holds its connection open and stops looking attractive. Long-lived connections such as uploads, WebSockets and server-sent events are the clearest case for it.
If you are unsure, least_conn is the safer default than round robin for an API with mixed endpoints. With uniform requests it behaves much like round robin, and with uneven ones it corrects itself. The cost is that nginx tracks connection counts per server.
Does IP hash give you sticky sessions, and what breaks?
ip_hash maps a client to a server by hashing its address. The nginx guide says it ensures requests from the same client always reach the same server except when that server is unavailable. The upstream module reference adds that the first three octets of an IPv4 address, or the entire IPv6 address, form the hashing key, and that a server being removed temporarily should be marked down to preserve the current hashing.
# Wrong for a pool that resizes or sits behind a proxy:
upstream api_pool {
ip_hash; # key = first three octets of the IPv4 address
server 10.0.0.11:3000;
server 10.0.0.12:3000;
server 10.0.0.13:3000 down; # "down" keeps the hashing of the others intact
}
# Modulo hashing, 12 keys, pool grows from 3 to 4 servers:
# key mod 3 == key mod 4 only for keys 0, 1, 2 -> 3 keep their server
# 9 of 12 keys (75 percent) move
# Right: hash a session identifier, with ketama consistent hashing.
# Requests with no cookie yet all hash the same empty key.
upstream api_sticky {
hash $cookie_sid consistent;
server 10.0.0.11:3000;
server 10.0.0.12:3000;
server 10.0.0.13:3000;
}
Two consequences follow from those details. First, hashing three octets means a whole /24 block, 256 addresses, maps to one server, so one office or mobile carrier network can pile onto a single replica. Second, plain modulo hashing remaps most clients when the pool size changes. With keys 0 to 11 and a pool growing from 3 to 4 servers, only keys 0, 1 and 2 keep the same server under key mod N, so 9 of 12 keys, 75 percent, move. The consistent parameter of hash uses ketama hashing, which the docs say remaps only a few keys when a server is added or removed.
Hashing the raw address is also the wrong key when something sits in front of nginx. Behind a CDN or another proxy every request arrives from the proxy address, so everyone hashes to the same server. Hash the thing that identifies the session instead, as below. Better still, keep session state in Redis so that no request needs a particular server at all; Wikipedia notes that persistence loses the session if its server fails, and that a shared database avoids that.
ip_hash does not make a pool resilient. Pinned clients stay pinned until nginx marks their server unavailable, and during a rolling deploy each restart drops the sessions held on that server unless they live in Redis or another shared store.
How do weights and power of two choices fit in?
Weights handle unequal machines and combine with the methods above. The nginx guide's example uses weight=3 on one server with two others at the default of 1, so out of every 5 new requests 3 go to the first server and 1 to each of the others. For a 4 vCPU node next to a 2 vCPU node, weight=2 against weight=1 gives a two thirds to one third split. Treat the weights as a starting guess, not a measurement.
# 3 of every 5 new requests go to the first server (3 + 1 + 1 = 5)
upstream api_weighted {
server 10.0.0.11:3000 weight=3;
server 10.0.0.12:3000;
server 10.0.0.13:3000;
}
# Power of two choices: pick two at random, keep the one with fewer
# active connections (least_conn is the default tie-break method).
upstream api_p2c {
random two;
server 10.0.0.11:3000;
server 10.0.0.12:3000;
server 10.0.0.13:3000;
}
Power of two choices is the middle route between random and least connections. Nginx exposes it as random two: it randomly selects two servers and then picks the one with fewer active connections by default. The HAProxy engineering blog explains the point: it saves the balancer from checking every server, which matters most when several balancers decide independently, since least connections on each could all choose the same quiet server. In HAProxy's own test under moderate contention, peak connection counts came out about 30 percent lower than with plain random or round robin, and least connections was about 4 percent better again. For one nginx in front of a handful of containers, plain least_conn is simpler and good enough.
Layer 4 or layer 7: does the algorithm change?
In an http block nginx balances layer 7: it has parsed the HTTP request, so it can hash a cookie, route by URL or terminate TLS before choosing. In a stream block it balances layer 4: it sees only a TCP connection. The stream upstream module offers hash, least_conn, least_time and random, and uses weighted round robin when none is named. It has no ip_hash directive, so the equivalent is hash $remote_addr, optionally with consistent.
# Layer 4: a stream block, TCP only. No ip_hash here; use hash on the address.
stream {
upstream pg_replicas {
least_conn; # long-lived connections: count them
server 10.0.0.21:5432;
server 10.0.0.22:5432;
}
server {
listen 5432;
proxy_pass pg_replicas;
}
}
# Layer 7 stays in the http block, where cookies, headers and URLs are visible.
Layer 4 fits protocols nginx should not parse, such as database traffic to read replicas. Connections there are long-lived, so counting connections with least_conn is more meaningful than alternating. Anything that needs the URL, a header or a cookie, including path-based routing, needs layer 7.
How do health checks change which algorithm works?
The algorithm only chooses among servers it believes are alive, so health checks decide that belief. Open-source nginx uses passive checks: when a response from a server fails, nginx marks it failed and avoids it for a while. The max_fails parameter defaults to 1 and 0 disables the accounting, and fail_timeout defaults to 10 seconds. After that time nginx probes the server with live client requests and restores it if they succeed.
upstream api_pool {
least_conn;
# After 3 failed attempts, skip the server for 30 s, then probe it
# with live requests. Defaults are max_fails=1 and fail_timeout=10s.
server 10.0.0.11:3000 max_fails=3 fail_timeout=30s;
server 10.0.0.12:3000 max_fails=3 fail_timeout=30s;
# Only receives traffic when the primary servers are unavailable.
server 10.0.0.13:3000 backup;
}
Passive checking costs a real user request per detection, so tighten it for production and keep a spare. The upstream module page lists the active health_check directive as part of the commercial subscription, so on the open-source build the passive settings plus a backup server are what you have. The algorithms react differently to a bad server: least_conn drifts away from a slow one by itself, while round robin keeps handing it its share until it is marked failed, and ip_hash keeps clients pinned to it until then. A proper /health endpoint on the app, as in the NestJS Terminus health checks post, is still worth having for deploys and monitoring.
Which load balancing algorithm should you choose?
Match the algorithm to what varies in your traffic. The table summarises the nginx directive, the situation that suits it and the way it fails.
Algorithm
nginx directive
Fits when
Breaks when
Round robin
none (default)
Identical stateless servers, uniform short requests
Request cost varies, so one server collects the slow ones
Weighted round robin
weight=N on server
Servers of different size
Weights are guesses that drift as the workload changes
Least connections
least_conn
Mixed request durations, long-lived connections
Connection count is a poor proxy for load when requests cost very differently per connection
IP hash
ip_hash
Quick stickiness with no shared session store
Clients share an address, a proxy sits in front, or the pool resizes
Hash with consistent
hash $cookie_sid consistent
Stickiness by session or cache key, pools that change size
The key is empty or skewed, so traffic concentrates
Power of two choices
random two
Several balancers deciding independently
A single balancer with a small pool, where least_conn is simpler
As a checklist, work down in this order and stop at the first line that applies:
Identical servers and similar request cost: round robin, which is already the default.
Servers of unequal size: add weight to the larger ones.
Request duration varies, or connections stay open for a long time: least_conn.
A client must return to the same server: move the session to Redis first. If that is impossible, hash a session identifier with consistent, not the raw IP.
In every case: set max_fails and fail_timeout, and keep a backup server.
On a single VPS running a few Docker replicas of a NestJS API, that checklist lands on least_conn with sessions in Redis, which is how I would set up an ERP or POS backend. Switch to something more elaborate only after the simple rule demonstrably misbehaves.
Choose the algorithm by what varies: server size, request cost, or the need for stickiness. Default to round robin for uniform work, least_conn for uneven work, and hashing only when state forces your hand. Whatever you pick, health checks decide whether it keeps working when a server fails.