Skip to main content

High Availability

Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/high-availability

In short

High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.

What is high availability?

High availability, often shortened to HA, describes a system designed to stay up and reachable nearly all the time. Availability is measured as the percentage of time a service works and is often described by its number of nines: 99.9% (three nines) allows about 8.8 hours of downtime per year, 99.99% about 53 minutes, and 99.999% (five nines) only about 5 minutes. Each extra nine is much harder and more expensive to achieve.

The core technique is removing single points of failure, meaning any component whose failure takes the whole system down. That means running several instances behind a load balancer, replicating databases, spreading servers across availability zones or regions, and using health checks to detect failed parts and move traffic away from them automatically, a process called failover. In an active-active setup all copies serve traffic at once, while in active-passive a standby takes over only when the primary fails, and release techniques such as rolling, blue-green, and canary deployments keep the service up during updates.

A hospital is a useful analogy: it has backup generators, several power feeds, and on-call staff so that care continues even when something breaks. Online stores, banks, payment systems, and communication services all aim for high availability because every minute of downtime costs money and trust, and their targets are usually written down as SLOs and promised to customers in service level agreements.

High availability is often confused with fault tolerance. A highly available system may have a short interruption, such as a few seconds of errors while traffic fails over, while a fault-tolerant system keeps working with no visible interruption at all. HA is also different from scalability, which is about handling more load, and from disaster recovery, which is about restoring service and data after a major event like the loss of a whole region.

Key takeaways

  • High availability means a system stays operational for a very high percentage of time.
  • Availability is often expressed in nines, such as 99.9% or 99.99%.
  • Redundancy, load balancing, replication, health checks, and failover remove single points of failure.
  • Every additional nine costs significantly more to achieve.
  • HA allows brief interruptions during failover; fault tolerance aims for none.

Example

Downtime budgets and the effect of redundancypython
# Allowed downtime per year for common availability targets
for target in [99.9, 99.99, 99.999]:
    minutes = (1 - target / 100) * 365 * 24 * 60
    print(f"{target}% -> {minutes:.0f} minutes per year")
# 99.9% -> 526, 99.99% -> 53, 99.999% -> 5

# Two services in series: a request fails if either one is down
print(0.999 * 0.999)  # 0.998001, about 99.8%

# Two redundant copies in parallel: down only if both fail at once
# (assuming their failures are independent)
print(1 - (1 - 0.999) ** 2)  # 0.999999, about 99.9999%

Readers ask

What does five nines mean?

Five nines means 99.999% availability, which allows only about 5 minutes of downtime per year. It is a very demanding target usually reserved for critical systems such as telecom networks and payment infrastructure.

What is the difference between high availability and fault tolerance?

High availability minimizes downtime but accepts a brief interruption while a backup takes over. Fault tolerance aims for no interruption at all, usually through fully redundant components running in parallel, which costs more.

What is a single point of failure?

A single point of failure is any component, such as one database server, one network link, or one person with the only access key, whose failure brings down the whole system. High availability design finds these components and duplicates them.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings