Short answers to what readers ask most about this topic.
01Should I scale horizontally or vertically first?
Scale vertically first for a single-service app, because it needs no code changes and no new moving parts. Move to horizontal scaling when you need to survive a machine failure, when resizing causes unacceptable downtime, or when the largest instance you can rent is no longer enough.
02What is the main disadvantage of vertical scaling?
It has a hard ceiling and a single point of failure. You can only buy as large an instance as your provider sells, and if that one machine fails, all of your capacity disappears at once. Resizing also often means a reboot or a migration.
03Why does horizontal scaling require a stateless app?
Once there are several instances, any request can land on any of them. If a session, cart or uploaded file lives only in one process's memory or disk, the next request will not find it. Moving that state into Redis, a database or object storage lets every instance serve every request.
04How do you handle sessions when you scale out?
Store sessions in a shared store with an expiry, such as Redis, so any instance can load them by id. Alternatively use signed tokens like JWT so the state travels with the request. Avoid sticky sessions, which the Twelve-Factor App says should never be relied upon.
05Why is the database the hardest part to scale horizontally?
Adding app servers is easy, but they all write to the same primary. PostgreSQL read replicas run read-only queries, so they add read capacity and never write capacity. Each new app instance also opens its own connection pool, which can exhaust max_connections, typically 100 by default.
Horizontal vs Vertical Scaling: Which to Pick First
Scale up first, scale out when you need redundancy. How statelessness, session handling, the database and cost shape decide between horizontal and vertical scaling.
Scale vertically first: resize the server until one machine stops being enough, because it needs no code changes. Scale horizontally when you need redundancy or more than one machine can give, which requires stateless app servers, sessions in a shared store like Redis, and a separate plan for the database, since writes still go to one primary.
Qilap, a carwash ERP I build, started life on one small VPS: NestJS, Postgres, Redis and Docker on the same box. The first time traffic looked like it might outgrow it, the question was the one every team asks: do I buy a bigger server, or add a second one?
This post answers that question in the order I would apply it. It uses the definitions and limits from the sources at the end (the Twelve-Factor App, the PostgreSQL manual, nginx and Amdahl's law), and every number below is arithmetic on stated assumptions, not a benchmark.
What is the difference between horizontal and vertical scaling?
Vertical scaling (scaling up) gives one machine more CPU, memory or disk. Horizontal scaling (scaling out) adds more machines and spreads work across them. The first changes the size of a node; the second changes the number of nodes.
That difference decides everything else. A bigger node is invisible to your code. More nodes are visible to everything: every request may now land on a different process, and anything your app remembered in memory is suddenly in the wrong place.
Which should you pick first?
Pick vertical first for a single-service app on one box, and move to horizontal when one of three things is true: you need to survive a machine failure, you cannot resize in place without downtime you can afford, or the biggest node you can rent is no longer enough. The table shows what each choice asks of you.
Question
Vertical (scale up)
Horizontal (scale out)
Code changes needed
None; the app does not know
App must be stateless, sessions and files moved out
If the machine dies
All capacity is gone at once
You lose one share of capacity, not all of it
Ceiling
The largest instance your provider sells
Far higher, bounded by shared components like the database
Resizing
Usually a reboot or a migration
Add or remove nodes behind a load balancer, no restart
Cost shape
Stepped: you buy the next size up
Smooth per node, plus a load balancer and spare capacity
Database
A bigger primary helps writes and reads
Replicas help reads only; writes stay on one primary
The bottom row is the one that surprises people. Adding app servers is the easy half of horizontal scaling. The data tier does not scale out the same way, which is why the database gets its own section below.
What does an app need before it can scale out?
It must be stateless: any request can be served by any instance, because nothing the user needs lives only in one process. The Twelve-Factor App states the rule as processes that are stateless and share-nothing, with anything that must persist stored in a backing service.
In-memory caches and counters that must agree across instances (a rate limiter, a cart, a queue of jobs).
Files written to local disk, such as uploads, generated PDFs and receipts, which the next instance cannot see.
Cron-style timers inside the app, which would now fire once per instance instead of once.
The fix is the same each time: move the state to a service every instance can reach. Here is the smallest version of the problem, with the broken and the working shape side by side.
// Wrong: this Map lives in ONE process. With two instances behind a
// load balancer, the cart written by instance A is invisible to instance B.
const carts = new Map<string, CartItem[]>();
app.post("/cart/items", (req, res) => {
const items = carts.get(req.userId) ?? [];
carts.set(req.userId, [...items, req.body]);
res.sendStatus(204);
});
// Right: the state lives in a service every instance can reach.
app.post("/cart/items", async (req, res) => {
const key = "cart:" + req.userId;
await redis.rpush(key, JSON.stringify(req.body));
await redis.expire(key, 60 * 60 * 24); // abandoned carts clean themselves up
res.sendStatus(204);
});
How do you handle user sessions across multiple servers?
Put the session in a shared store with an expiry, not in the server's memory. The Twelve-Factor App names Memcached or Redis as good candidates for session data because they offer time-expiration. Any instance can then load the session by id, so the load balancer is free to send a user's next request anywhere.
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL);
const SESSION_TTL_SECONDS = 1800; // 30 minutes, sliding
export async function saveSession(id: string, data: object) {
// SET key value EX seconds: the expiry is set atomically with the write
await redis.set("sess:" + id, JSON.stringify(data), "EX", SESSION_TTL_SECONDS);
}
export async function loadSession(id: string) {
const raw = await redis.get("sess:" + id);
if (!raw) return null;
await redis.expire("sess:" + id, SESSION_TTL_SECONDS); // slide the window
return JSON.parse(raw);
}
# nginx: round-robin is the default; least_conn favours the idle instance.
upstream api {
least_conn;
server 10.0.0.11:3000 max_fails=3 fail_timeout=10s;
server 10.0.0.12:3000 max_fails=3 fail_timeout=10s;
}
server {
listen 80;
location / {
proxy_pass http://api;
}
}
A signed token such as a JWT is the other route: the state travels with the request, so there is nothing to look up. The trade-off is that you cannot easily revoke it before it expires. I compare the two in an earlier post on JWT versus session authentication.
Sticky sessions pin a user to one instance, and they hide the problem rather than solve it. The Twelve-Factor App is blunt: they are a violation and should never be relied upon. When that instance restarts or is removed, every session it held is lost.
Why is the database the hard part of scaling?
Because every app instance you add still talks to the same primary. A PostgreSQL hot standby accepts only read-only queries while it replays the primary's write-ahead log, and the manual states that its transactions can never write. So replicas add read capacity; they do not add write capacity.
Amdahl's law puts a number on the ceiling. If a fraction s of the work is serial, the best possible speedup from n workers is 1 divided by (s plus (1 minus s) over n). Suppose 10 percent of each request is a serialised write to the primary. Eight app servers give 1 / (0.1 + 0.9/8) = 1 / 0.2125, about 4.7 times, and no number of servers can pass 10 times.
-- Amdahl: speedup = 1 / (s + (1 - s) / n), with s = 0.10 serial write share
-- n = 2 -> 1 / (0.10 + 0.45) = 1.82x
-- n = 8 -> 1 / (0.10 + 0.1125) = 4.71x
-- n = inf -> 1 / 0.10 = 10x (the ceiling)
-- Connection budget across the whole fleet, not per process:
-- instances * pool_size must stay below max_connections
-- 4 * 20 = 80 fits under the typical default of 100
-- 5 * 20 = 100 reaches it
SHOW max_connections;
SELECT count(*) FROM pg_stat_activity; -- how many are in use right now
The second trap is connections. The PostgreSQL manual says max_connections is typically 100 by default. If each instance opens a pool of 20, four instances already want 80 connections, and five reach the 100 limit. Count pools across the whole fleet, not per process, before you add nodes.
Before scaling out the app, ask whether the database is the actual bottleneck. If slow queries or missing indexes are the cause, a second app server doubles the load on the one component that was already struggling.
How does the cost differ between scaling up and scaling out?
Vertical cost is stepped; horizontal cost is smooth but carries overhead. Here is a worked example with assumed inputs, not measured ones. Say a request costs 40 ms of CPU, so one core serves 1000 / 40 = 25 requests per second. Planning for 70 percent utilisation gives 17.5 per core. A peak of 120 requests per second then needs 120 / 17.5 = 6.9, so 7 cores.
assumed CPU per request 40 ms -> 1000 / 40 = 25 req/s per core
target utilisation 70 % -> 25 * 0.7 = 17.5 req/s per core
peak load 120 req/s -> 120 / 17.5 = 6.86 -> 7 cores
1 x 8 vCPU 8 cores fails -> 0 cores -> 0 req/s
4 x 2 vCPU 8 cores lose 1 -> 6 cores -> 105 req/s (below 120)
5 x 2 vCPU 10 cores lose 1 -> 8 cores -> 140 req/s (survives)
price linear in vCPU (assumption): 10 units vs 8 -> 25 % premium for N+1
One 8 vCPU server covers that, but if it fails you serve nothing. Four 2 vCPU servers also total 8 cores, yet losing one leaves 6 cores, or 105 requests per second, which is short of the 120 peak. To survive one failure at peak you need five 2 vCPU servers: 10 cores, and 8 after a loss, or 140 per second. If price scales linearly with vCPU (check your provider; this is an assumption), that is 10 units against 8, a quarter more, plus the load balancer. That premium is what redundancy costs.
A checklist for deciding
Work through these in order. Stop at the first step that solves your problem, because each later step costs more engineering than the one before.
Measure first: find the actual bottleneck (CPU, memory, disk, or the database) before touching the topology.
Fix the cheap causes: indexes, N+1 queries, caching, connection pooling.
Scale up the single node if the bottleneck is CPU or memory and you can accept a short resize window.
Make the app stateless: sessions to Redis, files to object storage, scheduled jobs to a single runner.
Scale out behind a load balancer, add read replicas for read-heavy load, and size N+1 so one failure does not breach your peak.
For a small ERP or POS workload on one VPS, steps one to three are usually enough for a long time, and nothing in them closes the door on step four.
Scale up first because it is the cheapest change, and make your app stateless early because that is what keeps scaling out possible later. Treat the database as a separate problem: replicas add reads, never writes, and every new instance costs connections.