Fault Tolerance
- In Turkish
- Hata Toleransı
In short
Fault tolerance is the ability of a system to keep working correctly, perhaps at reduced capacity, when some of its hardware or software components fail.
What is fault tolerance?
Fault tolerance is the property that lets a system continue operating even when parts of it break. Disks fail, servers crash, networks drop packets, and dependencies time out; a fault-tolerant system expects these faults and is designed so they don't turn into failures that users notice. Engineers distinguish a fault, which is a defect or breakdown in one component, from a failure, which is when the system as a whole stops delivering its service.
The foundation is redundancy: extra copies of hardware, data, or services so another can take over when one breaks, as in RAID disk arrays, replicated databases, or clusters that use a majority vote, called a quorum, to agree on data. Software adds techniques such as timeouts, retries with exponential backoff, circuit breakers that stop calling a failing dependency, bulkheads that isolate resources so one failure can't consume everything, and idempotent operations that are safe to repeat. When something can't be recovered, graceful degradation keeps the core working, for example by showing cached recommendations instead of an error page.
A twin-engine airplane is the classic example: it is designed to fly and land safely on a single engine, so losing one doesn't cause a crash. Fault tolerance is essential in aviation, medical devices, databases, payment processing, and large distributed systems, where teams often test it deliberately with chaos engineering, injecting failures in a controlled way to confirm the system survives them.
Fault tolerance is often confused with high availability. High availability aims to minimize downtime and may accept a short interruption while traffic fails over to a backup, while fault tolerance aims for no interruption at all, which usually requires fully redundant components running in parallel and costs more. Fault tolerance is also broader than error handling in a single function: catching an exception is a local fix, while fault tolerance is a property of the whole system's design.
Key takeaways
- Fault tolerance keeps a system working correctly when components fail.
- A fault is a problem in one component; a failure is when the whole service stops working.
- Redundancy is the foundation, supported by timeouts, retries, and circuit breakers.
- Graceful degradation keeps core features working when others can't recover.
- Fault tolerance aims for no interruption; high availability accepts a brief one.
Example
// Try each replica in turn so one failed node doesn't fail the request
async function readWithFailover(replicas: string[], key: string) {
for (const url of replicas) {
try {
const res = await fetch(`${url}/items/${key}`, { signal: AbortSignal.timeout(2000) });
if (res.ok) return await res.json();
} catch {
// Timeout or network error: move on to the next replica
}
}
// Graceful degradation: every replica failed, so return a safe default
return { key, value: null, stale: true };
}Readers ask
What is the difference between fault tolerance and high availability?
Fault tolerance aims for a system that keeps working without interruption when a component fails, usually by running redundant components in parallel. High availability aims for minimal downtime and may allow a brief outage while a standby takes over.
What is the difference between a fault and a failure?
A fault is a problem in one part of a system, such as a crashed server or a corrupted disk. A failure is when the system as a whole stops providing its service, and fault tolerance tries to stop faults from becoming failures.
How do you test fault tolerance?
Teams use chaos engineering and failure-injection tests, deliberately killing servers, adding network latency, or blocking dependencies in a controlled way. They then check that the system keeps serving users and recovers as designed.
See also
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
- Circuit Breaker PatternSoftware Architecture, p. 4The circuit breaker pattern protects a system by stopping calls to a failing dependency for a while and failing fast instead of waiting on timeouts.
- Database ReplicationDatabases, p. 10Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
- Chaos EngineeringDevOps & Cloud, p. 8Chaos engineering is the practice of deliberately injecting failures into a system, such as crashing servers, to confirm that it keeps working as expected.
- Exponential BackoffBackend & APIs, p. 15Exponential backoff is a retry strategy that waits longer after each failed attempt, such as 1, 2, 4 and 8 seconds, so a struggling service can recover.
- IdempotencyBackend & APIs, p. 24Idempotency is the property of an operation that produces the same result whether it runs once or many times, so accidentally repeating a request is safe.
- Distributed SystemSoftware Architecture, p. 14A distributed system is a set of computers that work together over a network and appear to their users as a single system.
Spotted a mistake or something missing on this page?Suggest an edit