Short answers to what readers ask most about this topic.
01How do you design for high availability?
Find and remove single points of failure by running at least two of each tier, with health checks that detect a failed node and a mechanism that redirects traffic. Keep stateless tiers active-active with spare capacity, give the database a rehearsed promotion path, and test failover regularly. Then check the result against your downtime budget using the series and parallel availability formulas.
02How much downtime is 99.9% and 99.99% availability?
Over a 365-day year, 99.9% allows 8.76 hours of downtime and 99.99% allows 52.56 minutes. Per 30-day month that is 43.2 minutes and 4.32 minutes. The formula is one minus availability, multiplied by the length of the period.
03What is the difference between active-active and active-passive failover?
In active-active every node serves traffic and a failure shifts its share to the survivors, so they need spare capacity. In active-passive one node serves while a standby waits to be promoted, so failover includes detection and takeover time. Stateless app tiers suit active-active, and single-writer databases usually suit active-passive.
04How fast is failover with keepalived and VRRP?
RFC 5798 sets the backup's takeover timer to three times the advertisement interval plus a skew time. With a 1 second interval and backup priority 90 that is 3.648 seconds. A health script that needs several failed checks first adds to that, so a worst case near 10 seconds is realistic, but you should measure it in a failover test.
05Is DNS failover enough for high availability?
Usually not on its own, because resolvers cache a record for its TTL and a longer TTL delays changes taking effect. With a 60 second TTL and health checks every 10 seconds, clients can keep a dead address for around 90 seconds. DNS is the right tool for moving between sites, while a floating IP or a proxy is faster inside one site.
High Availability Design: Redundancy, Failover and the Nines
How to design for high availability: what each nine allows in minutes, series versus parallel availability maths, failover detection time and a keepalived and nginx sample.
To design for high availability, remove every single point of failure by running at least two of each tier behind a health-checked failover mechanism. Serial tiers multiply availability down, independent redundant pairs raise it to one minus the failure probability squared, detection time counts against the downtime budget, and the stateful database is the hard part.
The question I keep asking about single-VPS deployments, the kind a point-of-sale or ERP backend usually starts on, is plain: what happens at the busiest hour if that one machine is down? Not a disaster, just a kernel update that needs a reboot, a full disk, or a provider maintenance window.
This post answers how to design for high availability, from the arithmetic up. The nines table and the formulas are derived and shown, the calculator output is real output from my machine, and the config follows the keepalived and nginx documentation. I have not run this keepalived pair on production, so treat the config as a documented starting point to test, not a tested recipe.
What does 99.9% or 99.99% availability actually allow?
Availability is the fraction of time a service works, and it is usually quoted in nines, as the Google SRE book does with 99.9%, 99.99% and 99.999%. The conversion is one line: downtime equals one minus availability, times the length of the period. Over a 365-day year that is 8,760 hours, and the table below is that multiplication.
Availability
Downtime per year
Downtime per 30-day month
Failovers of 60 s per year
99%
87.60 hours
7.20 hours
5,256
99.5%
43.80 hours
3.60 hours
2,628
99.9%
8.76 hours
43.20 minutes
525
99.95%
4.38 hours
21.60 minutes
262
99.99%
52.56 minutes
4.32 minutes
52
99.999%
5.26 minutes
25.92 seconds
5
The last column is the one that changes designs. It divides the yearly budget by 60 seconds and rounds down: at 99.99% you can afford 52 one-minute failovers in a whole year, and at 99.999% only 5. The SRE book quotes the same 52.56 minutes for 99.99%. Once the target is that tight, a failover that takes minutes is itself the outage, and the design has to make detection and takeover fast.
How do redundancy and series dependencies change availability?
Tiers in a request path are in series: the proxy, the app and the database must all be up for a request to succeed, so their availabilities multiply. Three tiers at 99.9% each give 0.999 x 0.999 x 0.999 = 0.997003, which is 99.7003% or 26.25 hours of downtime a year. Adding a tier to a chain never makes it better, and the chain is always worse than its weakest link.
Redundant copies of one tier are in parallel: the tier is down only when every copy is down. For n independent copies of availability a, the tier availability is 1 minus (1 minus a) to the power n. A pair of 99.9% machines gives 1 minus 0.001 x 0.001 = 0.999999, so 99.9999% for that tier on paper.
Here is the calculator I ran, plain TypeScript with no dependencies, executed with a current Node that strips types. It also computes the VRRP takeover time I use in a later section.
const HOURS_PER_YEAR = 365 * 24; // 8760
// Every tier must be up, so availabilities multiply.
const serial = (tiers: number[]): number => tiers.reduce((acc, a) => acc * a, 1);
// n replicas, any one is enough. Assumes failures are independent.
const parallel = (a: number, n: number): number => 1 - Math.pow(1 - a, n);
const hoursDown = (a: number): number => (1 - a) * HOURS_PER_YEAR;
// RFC 5798: Master_Down_Interval = 3 * advert + (256 - priority) * advert / 256
const vrrpTakeover = (advertSec: number, backupPriority: number): number =>
3 * advertSec + ((256 - backupPriority) * advertSec) / 256;
const proxy = 0.999;
const app = 0.999;
const db = 0.999;
const single = serial([proxy, app, db]);
const everyTierPaired = serial([proxy, app, db].map((a) => parallel(a, 2)));
console.log("one of each: " + (single * 100).toFixed(4) + "% = " + hoursDown(single).toFixed(2) + " h/yr");
console.log("every tier paired: " + (everyTierPaired * 100).toFixed(4) + "% = " + (hoursDown(everyTierPaired) * 60).toFixed(2) + " min/yr");
console.log("99.99% budget: " + (hoursDown(0.9999) * 60).toFixed(2) + " min/yr");
console.log("VRRP takeover: " + vrrpTakeover(1, 90).toFixed(3) + " s (advert 1 s, backup priority 90)");
$ node ha-calc.ts
one of each: 99.7003% = 26.25 h/yr
every tier paired: 99.9997% = 1.58 min/yr
99.99% budget: 52.56 min/yr
VRRP takeover: 3.648 s (advert 1 s, backup priority 90)
Doubling only the app tier, which a longer version of the same script printed, lifts the stack to 99.8%, or 17.52 hours a year, because the proxy and database are still single and still in series. Pairing every tier gives 99.9997%, or 1.58 minutes a year. The lesson is that redundancy only pays where it removes the weakest remaining link, and that the last single box dominates the result.
The parallel formula assumes the two copies fail independently and that failover is instant and perfect. Two replicas on one host, one power feed, one deploy pipeline or one bad config push fail together, and then 1 minus (1 minus a) squared is fiction. Read the paired number as a ceiling, never as a forecast.
What is a single point of failure and how do you find one?
A single point of failure is any component whose failure stops the whole service because nothing can take over from it. The practical way to find them is to draw the request path and, for every box on it, ask what happens if it dies right now. Any answer of everything stops is a single point of failure. The usual suspects on a small stack are these.
The single VPS or virtual machine that hosts the proxy, the app and the database together.
The load balancer or reverse proxy itself, because redundant app servers behind one proxy only move the single point up a tier.
The one primary database or Redis instance that every app server reads and writes.
The one public IP address or DNS provider that all clients resolve through.
Shared things that fail every replica at once: an expiring TLS certificate, one secrets store, one deploy pipeline that ships the same bad build everywhere.
Fix them in order of how much of the request path each one takes down, not in order of how interesting they are. The proxy and the database usually come first, and the shared-fate items last, because they need process changes such as staged rollouts rather than extra servers.
Should you run active-active or active-passive?
In active-active every node serves traffic all the time, and failure means the survivors take a larger share. In active-passive one node serves and a standby waits, and failure means the standby has to be promoted. Neither is better in general, and the choice differs per tier.
Question
Active-active
Active-passive
Capacity you pay for
All of it serves traffic
The standby idles until needed
What failover involves
Remove the dead node from rotation, nothing to promote
Detect, then promote or move the address: the full detection budget applies
Main risk
Survivors overloaded by the lost share
The standby is untested or stale when you need it
Fits best
Stateless app tiers behind a proxy
A database or anything with a single writer
The overload risk is plain arithmetic. Two active nodes each running at 60% utilisation need 120% of one node when one fails, so the survivor falls over and the pair is down, not degraded. For n nodes, keep each below (n - 1) / n of capacity: 50% for a pair, about 67% for three. My default is active-active for stateless tiers and active-passive for the database.
How fast can failover detect a failure and take over?
Failover has three parts, detect, decide and redirect, and each spends downtime budget. Detection is usually the largest. The three common mechanisms, a floating IP with VRRP, a proxy that stops sending to a dead upstream, and DNS, differ mostly in how long each part takes.
For VRRP, RFC 5798 gives the backup's takeover timer: Master_Down_Interval equals 3 times the advertisement interval plus a skew time of (256 minus its priority) times the interval divided by 256. With an interval of 1 second and a backup priority of 90 that is 3 + 0.648 = 3.648 seconds, the figure my calculator printed. A health script in front of it adds more: with a check every 2 seconds and 3 failures needed, the master lowers its own priority after up to 6 seconds, and because a backup discards advertisements with a lower priority while preempt is on, which is the default, it then waits out its timer. That is roughly 6 + 3.65, close to 10 seconds in the worst case. This is my arithmetic from the documents, not a measurement.
# /etc/keepalived/keepalived.conf on node A (preferred). Node B is identical
# except: state BACKUP, priority 90.
vrrp_script chk_nginx {
script "/usr/bin/curl -fsS --max-time 1 http://127.0.0.1/healthz"
interval 2 # run the check every 2 s
fall 3 # 3 failed runs in a row = the check is KO
rise 2 # 2 good runs in a row = OK again
weight -20 # on KO, subtract 20 from priority: 100 -> 80, below B's 90
}
vrrp_instance VI_WEB {
state MASTER
interface eth0
virtual_router_id 51 # must match on both nodes
priority 100
advert_int 1 # advertise every 1 s
authentication {
auth_type PASS
auth_pass change-me
}
virtual_ipaddress {
203.0.113.10/24 # the address clients and DNS point at
}
track_script {
chk_nginx
}
}
That is what the keepalived config below encodes. The weight of minus 20 only works because it drops node A from 100 to 80, under node B's 90; a weight too small to cross that gap would never move the address. The nginx block covers the other layer: passive checks in the open-source upstream module, with active health checks being a commercial feature.
upstream app {
# Passive checks: 2 failures inside 10 s take a server out for the next 10 s.
server 10.0.0.11:3000 max_fails=2 fail_timeout=10s;
server 10.0.0.12:3000 max_fails=2 fail_timeout=10s;
# Only receives traffic when the two above are unavailable.
server 10.0.0.13:3000 backup;
}
server {
listen 80;
location = /healthz { return 200 "ok\n"; }
location / {
proxy_pass http://app;
# Retry the next server on connection errors, timeouts and 502/503.
# POST is NOT retried unless you add non_idempotent: leave it off for payments.
proxy_next_upstream error timeout http_502 http_503;
}
}
nginx does not pass a POST to the next server once it has been sent to an upstream, unless you add non_idempotent to proxy_next_upstream. For payments, leave it off. A retried charge without an idempotency key is a duplicate charge, and failover is exactly when retries happen.
DNS failover is the slow option and the TTL is why. The Route 53 documentation describes TTL as how long resolvers cache a record, and warns that a longer value delays changes taking effect. With a 60 second TTL, health checks every 10 seconds and an assumed three failures needed, a client can keep the old address for 30 + 60 = 90 seconds. At 99.99% that is 3,153.6 seconds of yearly budget divided by 90, so only 35 such failovers, against over 320 for the roughly 9.65 second VRRP path. Use DNS where you must cross sites, and a floating IP or proxy inside one.
Why is the stateful tier the hard part?
Stateless tiers are redundant by copying. Start another container and it is identical. A database cannot be copied that way, because the copies have to agree on every write. Each extra copy therefore costs a consistency decision, and a failover has to decide which copy is the truth.
The detail that bites is the commit mode. With asynchronous replication the primary acknowledges a write before a replica has it, so a failover can lose the newest committed writes. With synchronous replication it waits, which costs latency and, if the replica is down, availability too. I covered the mechanics in my post on leader-follower replication, and the lag pitfalls of reading from replicas in the read replica post. For a point-of-sale system, a lost order after a promote is a cash-drawer mismatch, so the trade matters.
Automatic promotion also adds split-brain risk: if the old primary is only partitioned, not dead, two nodes accept writes. Fencing the old primary before promoting is what prevents it. My rule for a small team is that a rehearsed, documented manual promote beats automation nobody has ever triggered, and that sessions, uploads and caches should be moved out of the app tier first so that any node can serve any request.
Is multi-AZ enough, or do you need multi-region?
An AWS Availability Zone is one or more discrete data centres with separate power, networking and connectivity, and AWS describes zones in a region as up to about 100 km apart: far enough to avoid correlated failures, close enough for synchronous replication at single-digit millisecond latency. That combination is why multi-AZ is the default answer, because you get independent failure domains without giving up synchronous replication.
Multi-region protects against a whole region failing, but the distance makes replication effectively asynchronous, doubles the infrastructure and brings the data-loss window back. Start with multi-AZ, and go multi-region only when a target requires it. On a plain VPS the equivalent check is whether your two servers are really in separate facilities or only separate machines on the same rack, a point my disaster recovery post for a VPS also raises.
How do you test failover before it matters?
A failover that has never run is a hypothesis. A game day is a scheduled, announced exercise where you break something on purpose and compare what happens with what you predicted. A minimal sequence for the setup above looks like this.
Write the prediction first: the virtual IP moves within 10 seconds and no request fails after that.
Stop the nginx process on the master and measure the time until the backup answers on the virtual IP.
Reboot the whole master node, which tests the case a service restart hides.
Break a dependency, such as blocking the database port from one app node, and watch the health check and the retries.
Fail back, then confirm the old master rejoined as a standby and that no data was lost.
Record the measured time against the budget from the first section, fix the gap and repeat. Rolling out deploys one node at a time, as in my zero-downtime deploy posts, quietly exercises the same path every release, which is why I like it as a standing test.
Do the arithmetic before buying anything. Remove the single points of failure in the order of how much each one takes down, run stateless tiers active-active with headroom, give the database a rehearsed promote, and spend the downtime budget on detection time you measured in a game day, not on the figure from a datasheet.