Consensus Algorithm
- In Turkish
- Konsensüs Algoritması
- Pronunciation
- kun-SEN-sus AL-guh-rith-um
In short
A consensus algorithm lets a group of machines agree on a single value or an ordered log of decisions, even when some of them crash or messages are lost.
What is a consensus algorithm?
Replicating data across several servers only helps if they agree on what the data is. If two servers both think they are in charge, or accept different writes in a different order, the copies drift apart. Consensus algorithms solve this: the group elects a leader, the leader proposes entries, and an entry counts as committed only once a majority, called a quorum, has stored it.
The two best-known algorithms are Paxos, described by Leslie Lamport and published in 1998, and Raft, published in 2014 and designed to be easier to understand and implement. Raft is used in etcd, which stores Kubernetes' cluster state, in Consul and in CockroachDB, and Kafka's KRaft mode uses a Raft-based protocol to manage its own metadata. ZooKeeper uses a similar protocol called ZAB.
Majorities set the fault tolerance. A cluster of 3 nodes keeps working if 1 fails, and a cluster of 5 survives 2 failures, which is why clusters have an odd number of members. If a network split leaves no side with a majority, the system stops accepting writes rather than risk two conflicting histories, choosing consistency in the CAP sense.
A common misconception is that every distributed database runs consensus for every write. Consensus is relatively slow, because each decision needs a round trip to a majority, so many systems use it only for coordination, such as electing a leader or storing configuration. Blockchains face a harder problem, Byzantine fault tolerance, where some participants may lie, and use different mechanisms such as proof of stake.
Key takeaways
- Consensus lets machines agree on values or an ordered log despite failures.
- Raft and Paxos are the best-known algorithms; Raft is easier to implement.
- A leader proposes entries, which commit once a majority stores them.
- 3 nodes tolerate 1 failure; 5 nodes tolerate 2.
- etcd, Consul and Kafka's KRaft rely on Raft for coordination.
Readers ask
What is the difference between Raft and Paxos?
Both solve the same problem with similar guarantees. Paxos came first and is famously hard to understand and implement completely. Raft was designed for clarity, with a strong leader and clearly separated steps, and is now the more common choice.
Why do clusters use an odd number of nodes?
Because consensus needs a majority. Four nodes tolerate only one failure, the same as three, so adding the fourth costs more without adding fault tolerance. Odd sizes such as 3, 5 or 7 make the most of each node.
What is leader election?
The part of a consensus protocol that chooses which node coordinates the group. If the leader fails or becomes unreachable, the remaining nodes hold an election and pick a new one.
See also
- Distributed SystemSoftware Architecture, p. 14A distributed system is a set of computers that work together over a network and appear to their users as a single system.
- Database ReplicationDatabases, p. 10Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
- CAP TheoremDatabases, p. 2The CAP theorem says that if a network failure splits a distributed database, the system must choose between consistency and availability; it can't have both.
- Fault ToleranceSoftware Architecture, p. 20Fault tolerance is the ability of a system to keep working correctly, perhaps at reduced capacity, when some of its hardware or software components fail.
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
- KafkaBackend & APIs, p. 26Apache Kafka is a distributed event streaming platform that stores events in durable, ordered logs for many services to publish, read in real time or replay.
Spotted a mistake or something missing on this page?Suggest an edit