SLA
Service Level Agreement
- Pronunciation
- es-el-AY
In short
An SLA (service level agreement) is a provider's commitment to customers about the level of service, such as 99.9% uptime, and what happens if it isn't met.
What is an SLA?
An SLA turns reliability into a promise. A cloud database might guarantee 99.95% monthly availability, a support plan might promise a first response within an hour, and an API might commit to a maximum error rate. If the provider falls short, the agreement says what the customer gets, typically a credit off the bill.
The number of nines matters more than it looks. 99.9% availability allows about 43 minutes of downtime in a 30-day month, while 99.99% allows only about 4.3 minutes. Each extra nine requires more redundancy, automation and on-call effort, so higher SLAs cost much more to provide.
SLAs sit on top of two internal ideas from site reliability engineering. A service level indicator (SLI) is the measurement, such as the share of successful requests; a service level objective (SLO) is the internal target for it. Teams set their SLOs stricter than the SLA, so they get warned and can act before a contractual promise is broken.
A common misconception is that a service's availability is the same as its SLA. An application built on several services multiplies their risks: if it depends on three components that are each 99.9% available, its own availability is lower than any one of them, unless it is designed to tolerate their failures.
Key takeaways
- An SLA is a formal promise about service quality, such as uptime.
- Missing it usually earns customers service credits.
- 99.9% allows about 43 minutes of downtime a month; 99.99% about 4.3 minutes.
- SLIs measure, SLOs set internal targets, and SLAs make external promises.
- Dependencies combine, so a system can be less available than its parts.
Readers ask
What is the difference between an SLA, an SLO and an SLI?
An SLI is a measurement, such as the percentage of successful requests. An SLO is the target a team sets for it. An SLA is the agreement with customers, with consequences if the promised level isn't met.
What does 99.9% uptime mean?
The service is available at least 99.9% of the time in the measured period, which allows about 43 minutes of downtime in a 30-day month, or nearly 9 hours over a year.
What happens when an SLA is breached?
Usually the customer can claim a service credit, a percentage of the monthly fee depending on how far availability fell. Serious or repeated breaches may allow the customer to end the contract.
Often compared
See also
- SLODevOps & Cloud, p. 51An SLO is a measurable reliability target for a service, such as 99.9% of requests succeeding over 30 days, that tells a team how reliable is reliable enough.
- Site Reliability EngineeringDevOps & Cloud, p. 49Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- PostmortemDevOps & Cloud, p. 42A postmortem is a written review after an incident that explains what happened, why it happened, and what the team will change so it doesn't happen again.
- SaaSDevOps & Cloud, p. 46SaaS (software as a service) is software delivered over the internet as a subscription; users sign in while the provider runs, updates and secures it.
Spotted a mistake or something missing on this page?Suggest an edit