Short answers to what readers ask most about this topic.
01How does the Raft consensus algorithm work?
Raft elects one leader per term. The leader appends client commands to its log, replicates them to followers, and marks an entry committed once a majority stores it. If the leader fails, a follower whose election timeout fires first becomes a candidate and wins with a majority of votes.
02Why does Raft use randomized election timeouts?
If every follower used the same timeout they would all become candidates at once and split the votes repeatedly. A random timeout, such as 150 to 300 ms in the paper's example, means one follower usually fires first and wins before the others start. If a split vote still happens, each candidate draws a new random timeout for the retry.
03How many node failures can a Raft cluster tolerate?
A cluster of n servers needs floor(n / 2) + 1 of them to form a quorum, so it tolerates n minus that many failures. That gives 1 failure for 3 nodes, 2 for 5 nodes and 3 for 7 nodes. Four nodes still tolerate only 1, which is why odd sizes are recommended.
04What is the difference between a committed and an applied entry in Raft?
An entry is committed when the leader knows it is stored on a majority of servers. Each server then applies committed entries to its state machine in log order. The leader shares its commit index through later AppendEntries messages, so followers apply an entry slightly after the leader does.
05Which systems use Raft?
etcd uses Raft for its replicated key-value store, and its FAQ describes a leader-based design that persists the Raft log to disk. HashiCorp Consul runs Raft among its server agents and documents quorum sizes for 3 to 8 servers. The Raft website lists many other implementations.
Raft Consensus Algorithm Explained: Elections and Logs
How the Raft consensus algorithm works: terms and roles, randomized leader election, log replication and the commit index, quorum arithmetic for 3, 5 and 7 nodes, safety, and where etcd and Consul use it.
Raft keeps a cluster agreeing on one ordered log by electing a single leader per term with randomized timeouts, replicating each entry to a majority before committing it, and refusing votes to candidates with stale logs. A majority quorum means 3 nodes tolerate 1 failure, 5 tolerate 2, and 7 tolerate 3.
Once a stateful service runs on more than one node, the question that matters is not how to copy data but who is allowed to decide. My own systems, an ERP, a POS and a carwash back office on NestJS and Postgres, let the database own that decision, so I never had to build the answer myself. Raft is the answer that the databases and coordination services underneath use.
This post answers one question: how does the Raft consensus algorithm work? It follows the original paper by Diego Ongaro and John Ousterhout, derives the quorum numbers instead of quoting them, and ends with a small TypeScript toy that I ran, so the output shown is real. The toy is for understanding. It is not an implementation, and the post says exactly what it leaves out.
Why do distributed systems need consensus?
Consensus is how several servers agree on one sequence of commands even when some of them crash or lose the network. The paper frames it through the replicated state machine: every server holds a copy of the same state machine and applies the same commands in the same order, so every copy ends in the same state. The consensus algorithm manages the replicated log of those commands. If the log is identical everywhere, the state is identical everywhere.
The hard part is that a server that stops answering is indistinguishable from a server that is slow, and two halves of a split network can each believe they are in charge. Raft solves this by making one server the leader, letting only the leader order commands, and requiring a majority to accept each command before it counts. As long as a majority of servers can talk to each other, the cluster keeps accepting writes. Below a majority it stops, which is a deliberate choice of consistency over availability.
What are terms and roles (follower, candidate, leader) in Raft?
Every Raft server is in exactly one of three roles at any moment, and time is cut into numbered terms. A term is a logical clock that only moves forward: it starts with an election and lasts until the next one. Servers exchange term numbers in every message, and any server that sees a higher term than its own immediately adopts it and steps down to follower.
Role
What it does
How it changes role
Follower
Passive. Answers requests from leaders and candidates and applies committed entries.
Becomes a candidate when the election timeout elapses with no valid heartbeat.
Candidate
Increments its term, votes for itself and asks the others for votes.
Wins with a majority and becomes leader, or sees a leader or higher term and steps down, or times out and retries.
Leader
Handles every client write, replicates it, and sends heartbeats to hold its authority.
Steps down to follower if it sees a higher term.
The term number is what makes stale leaders harmless. A leader that was cut off by the network and comes back still believes in its old term. The first message it sends or receives carries a higher term, so it steps down instead of overwriting anything. Each server also votes for at most one candidate per term, so a term can produce at most one leader.
How does leader election work with randomized timeouts?
Servers start as followers. The leader sends periodic heartbeats, which are AppendEntries messages with no entries in them. A follower that hears nothing for a period called the election timeout assumes there is no viable leader, increments its term, becomes a candidate, votes for itself and sends RequestVote to everyone else. The paper's example is a timeout drawn at random from an interval such as 150 to 300 ms.
broadcastTime << electionTimeout << MTBF (Raft paper, section 9.3)
heartbeat interval : well under the timeout (heartbeats are empty AppendEntries)
election timeout : drawn at random per node, per attempt, e.g. 150-300 ms
split vote : nobody reaches quorum -> every candidate waits a NEW random
timeout -> one of them usually fires first and wins the retry
The randomness is the whole trick. If every follower used the same timeout, they would all become candidates together, split the votes and fail to reach a majority, over and over. With random timeouts one follower usually fires first, wins its election and starts sending heartbeats before the others wake up. If a split vote does happen, each candidate picks a new random timeout for the retry. The paper also gives the timing constraint that makes this stable: the broadcast time should be an order of magnitude below the election timeout, and the election timeout a few orders of magnitude below the mean time between failures.
The election timeout is also your failover time. While the old leader is gone, the cluster accepts no writes for roughly one timeout. Tune it above your real network round trip and disk sync latency, otherwise healthy followers will start needless elections.
How does log replication and the commit index work?
A client sends a command to the leader. The leader appends it to its own log as a new entry stamped with the current term, then sends AppendEntries to every follower. Each message names the entry just before the new ones, by index and term. A follower accepts only if its log has a matching entry at that position. This consistency check is how Raft guarantees the Log Matching property: if two logs contain an entry with the same index and term, the logs are identical up to that entry.
index: 1 2 3 4
term: 1 1 2 2
cmd: SET s=10 DEC s 3 DEC s 2 DEC s 1
A replicated state machine applies the SAME log in the SAME order on every node:
s = 10 -> 7 -> 5 -> 4 (every node ends at s = 4)
committed = replicated on a majority AND safe to apply
commitIndex = 3 means entries 1..3 may be applied; entry 4 is not safe yet
AppendEntries(term, prevLogIndex, prevLogTerm, entries[], leaderCommit)
follower rejects unless its log has an entry at prevLogIndex with term prevLogTerm
(this is the Log Matching check; on rejection the leader backs nextIndex up by one)
An entry is committed once the leader knows it is stored on a majority of servers. The leader tracks this as the commit index and includes it in later AppendEntries messages, which is how followers learn what is safe to apply to their state machines. If a follower's log has diverged, the leader walks its nextIndex for that follower backwards until the logs agree and then overwrites the follower's conflicting entries. Followers never overwrite the leader.
How many failures can 3, 5 and 7 nodes tolerate?
A quorum is a majority, and the arithmetic for it fits in four lines. A cluster of n servers needs floor(n / 2) + 1 of them to agree, so it survives n minus that many failures. The deeper reason a majority is right is that any two majorities must share at least one server, and that shared server is what carries the knowledge of the last committed entry into the next leader.
quorum(n) = floor(n / 2) + 1
tolerated(n) = n - quorum(n)
n = 3: quorum = floor(3/2) + 1 = 2 tolerated = 3 - 2 = 1
n = 5: quorum = floor(5/2) + 1 = 3 tolerated = 5 - 3 = 2
n = 7: quorum = floor(7/2) + 1 = 4 tolerated = 7 - 4 = 3
n = 4: quorum = floor(4/2) + 1 = 3 tolerated = 4 - 3 = 1 (same as n = 3)
Why a majority is the right size: two quorums of size q out of n nodes
overlap in at least 2q - n nodes. With q = floor(n/2) + 1 that is always >= 1.
n = 5, q = 3: 3 + 3 - 5 = 1 node is in BOTH quorums
That shared node is the witness: it saw the last committed entry, so a new
leader that needs its vote cannot be elected with a log that lacks it.
Nodes (n)
Quorum
Failures tolerated
Note
3
2
1
The smallest cluster that survives a failure.
4
3
1
Same tolerance as 3 nodes, with a larger quorum to reach.
5
3
2
The common production choice for higher availability.
7
4
3
Survives more, but every write waits for 4 acknowledgements.
The same table appears in the Consul documentation, which recommends 3 or 5 servers for production and warns against a single server outside development. Notice what a bigger cluster does not buy: it does not make writes faster. Every commit still has to reach a majority through one leader, so adding nodes trades throughput and latency for fault tolerance.
Even cluster sizes are a trap. Four nodes tolerate one failure, exactly like three, but need three acknowledgements per write instead of two. Two nodes tolerate none, because losing either leaves one out of two, which is not a majority. Run an odd number.
What keeps Raft safe? The election restriction
Safety rests on one rule: a candidate cannot win unless its log is at least as up to date as the log of every server in the majority that votes for it. Up to date means a higher last term wins, and with equal last terms the longer log wins. Because a committed entry lives on a majority, and the winning candidate needs votes from a majority, at least one voter holds every committed entry and will refuse a candidate that lacks it. This is how the paper gets the Leader Completeness property: a leader's log always contains every entry committed in earlier terms.
There is one more subtlety. A leader may only commit entries from its own term by counting replicas. Older entries become committed indirectly, once a newer entry on top of them is. The paper's Figure 8 shows how counting replicas of an old entry can let a committed-looking entry be overwritten. The toy below enforces both rules and I ran it for real, with node2 cut off from the first write.
// raft-toy.ts (excerpt): the three rules that carry the safety argument.
// Teaching toy: single process, "RPCs" are method calls, no disk, no snapshots,
// no membership change, no real timers. NOT an implementation of Raft.
// RequestVote receiver: term rule, one vote per term, election restriction.
requestVote(term: number, cand: number, lastIdx: number, lastTerm: number): boolean {
if (term > this.term) { this.term = term; this.votedFor = null; this.role = "follower"; }
if (term < this.term) return false;
const upToDate =
lastTerm > this.lastTerm() ||
(lastTerm === this.lastTerm() && lastIdx >= this.log.length);
if (!upToDate) return false; // never vote for a log older than ours
if (this.votedFor !== null && this.votedFor !== cand) return false;
this.votedFor = cand;
return true;
}
// AppendEntries receiver: consistency check on prevIdx/prevTerm, then append.
appendEntries(term: number, prevIdx: number, prevTerm: number,
entries: Entry[], leaderCommit: number): boolean {
if (term < this.term) return false;
this.term = term; this.role = "follower"; // a valid leader exists this term
if (prevIdx > this.log.length ||
(prevIdx > 0 && this.log[prevIdx - 1].term !== prevTerm)) return false;
this.log = this.log.slice(0, prevIdx).concat(entries);
this.commitIndex = Math.max(this.commitIndex, Math.min(leaderCommit, this.log.length));
return true;
}
// Leader commit rule: majority AND an entry of the CURRENT term (Figure 8).
if (stored >= quorum && this.log[idx - 1].term === this.term) this.commitIndex = idx;
// Driver: 3 nodes, timeouts drawn from 150-300 ticks with a seeded PRNG (seed 7).
// 1) tick until the first timeout fires -> that node runs an election
// 2) leader.propose("SET stock=7", [node1]) // node2 is unreachable
// 3) leader crashes; stale node2 starts an election first
// 4) node1 (up to date) starts an election
$ node raft-toy.ts
1) boot: all followers, timeouts = 151, 159, 296 ticks
t=151 node0 starts election for term 1, votes=3/3, quorum=2
node0 is leader of term 1
node0 leader term=1 log=[] commit=0
node1 follower term=1 log=[] commit=0
node2 follower term=1 log=[] commit=0
2) client write while node2 is partitioned away (only node1 reachable)
node0 appended "SET stock=7" at index 1, stored on 2/3, commitIndex=1
node0 leader term=1 log=["1:SET stock=7"] commit=1
node1 follower term=1 log=["1:SET stock=7"] commit=0
node2 follower term=1 log=[] commit=0
3) leader crashes, partition heals; the STALE node (no entry) asks for votes first
t=152 node2 starts election for term 2, votes=1/3, quorum=2
stale node became leader? false
4) the up-to-date node runs its own election and wins with the stale node's vote
t=153 node1 starts election for term 3, votes=2/3, quorum=2
node1 is leader of term 3
node0 DOWN term=1 log=["1:SET stock=7"] commit=1
node1 leader term=3 log=["1:SET stock=7"] commit=0
node2 follower term=3 log=[] commit=0
Read the output closely. Node2 holds no entry, so when it asks for votes in term 2 it gets only its own, 1 of 3, and cannot win even though the old leader is down. Node1, which has the committed entry, then wins term 3 with node2's vote. The log survived the leader's death. Two honest details are visible too: the followers show commit equal to 0 because the commit index travels on the next AppendEntries, and the new leader shows 0 because a fresh leader cannot count old entries. The paper has each new leader commit a blank no-op entry at the start of its term for exactly this reason.
What the toy leaves out is most of Raft: persistence of term, vote and log before replying, real timers and heartbeats, message loss and reordering, the backwards walk of nextIndex, snapshots and log compaction, and membership changes. It teaches the rules. It is not something to deploy.
How do you change cluster membership safely?
Swapping servers in and out is dangerous because different servers can switch from the old configuration to the new one at different moments, and for a window two disjoint majorities could each elect a leader. The Raft paper avoids this with a two phase approach called joint consensus. The cluster first moves to a transitional configuration that combines old and new, where decisions need majorities of both. Once that is committed, it moves to the new configuration alone.
In practice you rarely implement this yourself, and you should change one server at a time. Check what your tool supports. The etcd FAQ, for example, says a node that fails permanently can be removed from the cluster through runtime reconfiguration.
Where is Raft used, and do you need to run it yourself?
Raft sits under real infrastructure. The etcd FAQ states that it is leader based, that a request needing consensus sent to a follower is forwarded to the leader, and that etcd persists the Raft log to disk and replays it after a power loss. Consul's documentation describes its servers as a Raft peer set where a quorum must agree before a state change is committed. The Raft site lists many more implementations.
For a typical ERP or POS stack on one VPS, you will not write Raft. A single Postgres primary or a Redis instance is simpler, and the consensus lives inside the managed service or the coordination layer you adopt. Use this checklist before reaching for it.
Do you need several nodes to agree on one ordered history of writes, not just copy data one way?
Can the service tolerate refusing writes when a majority is unreachable? Raft chooses consistency over availability.
Is there an existing system (etcd, Consul, a database with built in consensus) that already provides it, so you do not implement the protocol?
Will you run an odd number of servers, three for one failure or five for two, in separate failure domains?
Have you set the election timeout above your measured network and disk latency, and planned how members are replaced?
Raft earns its reputation for being understandable by reducing consensus to three small ideas: one leader per term chosen with randomized timeouts, a log that followers must match before they append, and a majority rule backed by the election restriction. Remember the arithmetic, floor(n / 2) + 1, odd cluster sizes, and treat the protocol as something you adopt through a mature system rather than write.