Skip to main content

Fault Tolerance

In Turkish
Hata Toleransı
Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/fault-tolerance

In short

Fault tolerance is the ability of a system to keep working correctly, perhaps at reduced capacity, when some of its hardware or software components fail.

What is fault tolerance?

Fault tolerance is the property that lets a system continue operating even when parts of it break. Disks fail, servers crash, networks drop packets, and dependencies time out; a fault-tolerant system expects these faults and is designed so they don't turn into failures that users notice. Engineers distinguish a fault, which is a defect or breakdown in one component, from a failure, which is when the system as a whole stops delivering its service.

The foundation is redundancy: extra copies of hardware, data, or services so another can take over when one breaks, as in RAID disk arrays, replicated databases, or clusters that use a majority vote, called a quorum, to agree on data. Software adds techniques such as timeouts, retries with exponential backoff, circuit breakers that stop calling a failing dependency, bulkheads that isolate resources so one failure can't consume everything, and idempotent operations that are safe to repeat. When something can't be recovered, graceful degradation keeps the core working, for example by showing cached recommendations instead of an error page.

A twin-engine airplane is the classic example: it is designed to fly and land safely on a single engine, so losing one doesn't cause a crash. Fault tolerance is essential in aviation, medical devices, databases, payment processing, and large distributed systems, where teams often test it deliberately with chaos engineering, injecting failures in a controlled way to confirm the system survives them.

Fault tolerance is often confused with high availability. High availability aims to minimize downtime and may accept a short interruption while traffic fails over to a backup, while fault tolerance aims for no interruption at all, which usually requires fully redundant components running in parallel and costs more. Fault tolerance is also broader than error handling in a single function: catching an exception is a local fix, while fault tolerance is a property of the whole system's design.

Key takeaways

  • Fault tolerance keeps a system working correctly when components fail.
  • A fault is a problem in one component; a failure is when the whole service stops working.
  • Redundancy is the foundation, supported by timeouts, retries, and circuit breakers.
  • Graceful degradation keeps core features working when others can't recover.
  • Fault tolerance aims for no interruption; high availability accepts a brief one.

Example

Reading from replicas with failover and a safe fallbacktypescript
// Try each replica in turn so one failed node doesn't fail the request
async function readWithFailover(replicas: string[], key: string) {
  for (const url of replicas) {
    try {
      const res = await fetch(`${url}/items/${key}`, { signal: AbortSignal.timeout(2000) });
      if (res.ok) return await res.json();
    } catch {
      // Timeout or network error: move on to the next replica
    }
  }
  // Graceful degradation: every replica failed, so return a safe default
  return { key, value: null, stale: true };
}

Readers ask

What is the difference between fault tolerance and high availability?

Fault tolerance aims for a system that keeps working without interruption when a component fails, usually by running redundant components in parallel. High availability aims for minimal downtime and may allow a brief outage while a standby takes over.

What is the difference between a fault and a failure?

A fault is a problem in one part of a system, such as a crashed server or a corrupted disk. A failure is when the system as a whole stops providing its service, and fault tolerance tries to stop faults from becoming failures.

How do you test fault tolerance?

Teams use chaos engineering and failure-injection tests, deliberately killing servers, adding network latency, or blocking dependencies in a controlled way. They then check that the system keeps serving users and recovers as designed.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings