SLO
Service Level Objective
In short
An SLO is a measurable reliability target for a service, such as 99.9% of requests succeeding over 30 days, that tells a team how reliable is reliable enough.
What is an SLO?
A service level objective, or SLO, is an internal target for how well a service should perform, expressed as a number over a period of time. A typical SLO reads: 99.9% of checkout requests will succeed within 300 milliseconds, measured over a rolling 30 days. SLOs come from site reliability engineering and give teams a shared, objective definition of good enough reliability.
Every SLO is built on a service level indicator (SLI), the actual measurement, such as the ratio of successful requests to all requests or the share of responses faster than a threshold. The SLO sets the target for that indicator, and the remaining gap to 100% is the error budget: with a 99.9% target, the service may fail 0.1% of the time, which is about 43 minutes of downtime in 30 days. Teams get alerted when the budget is burning too fast, and if it runs out, they pause risky releases and invest in reliability work.
An SLO works like a monthly spending budget: you don't aim to spend nothing, you agree on a limit and adjust when you get close to it. Aiming for 100% is almost never right, because each extra nine costs far more effort and users can't tell the difference, especially when their own networks and devices fail more often than that. SLOs are set for user-facing journeys such as signing in, searching, or paying, and they turn vague complaints like the site feels slow into numbers a team can track.
SLOs are often confused with SLAs. A service level agreement (SLA) is a contract with customers that promises a level of service and specifies penalties, such as refunds, if it is missed, while an SLO is an internal goal. SLOs are usually stricter than any SLA, so the team notices and fixes problems before a contractual promise is broken.
Key takeaways
- An SLO is a target value for a reliability measurement over a time window.
- The measurement itself is called an SLI, for example the share of successful requests.
- The allowed failure, 100% minus the SLO, is the error budget.
- An SLA is an external contract with penalties; an SLO is an internal goal and usually stricter.
- Targets of 100% are unrealistic and needlessly expensive.
Example
# SLO: 99.9% of requests succeed over a 30-day window
slo = 0.999
minutes_in_window = 30 * 24 * 60 # 43,200 minutes
error_budget = (1 - slo) * minutes_in_window
print(f"Allowed downtime: {error_budget:.1f} minutes") # 43.2 minutes
# 30 minutes of outages have already happened this month
remaining = error_budget - 30
print(f"Budget left: {remaining:.1f} minutes") # 13.2 minutes
if remaining < 0.5 * error_budget: # more than half is already spent
print("Freeze risky releases and focus on reliability")Readers ask
What is the difference between SLI, SLO, and SLA?
An SLI is the measurement, such as the percentage of successful requests. An SLO is the target for that measurement, such as 99.9% over 30 days, and an SLA is a contract with customers that promises a level of service and defines penalties if it is not met.
What is an error budget?
An error budget is the amount of unreliability an SLO allows, calculated as 100% minus the target. A 99.9% monthly availability SLO gives a budget of about 43 minutes of downtime, which teams can spend on risky releases and experiments.
How many nines should an SLO have?
It depends on what users need and on the reliability of the services it depends on. Many web services use 99.9% or 99.95%, while 99.99% allows only about 4 minutes of downtime per month and requires much more investment in redundancy and automation.
Often compared
See also
- Site Reliability EngineeringDevOps & Cloud, p. 49Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- MetricsDevOps & Cloud, p. 36Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
- PostmortemDevOps & Cloud, p. 42A postmortem is a written review after an incident that explains what happened, why it happened, and what the team will change so it doesn't happen again.
- LatencyNetworking, p. 14Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.
- SLADevOps & Cloud, p. 50An SLA (service level agreement) is a provider's commitment to customers about the level of service, such as 99.9% uptime, and what happens if it isn't met.
Spotted a mistake or something missing on this page?Suggest an edit