Skip to main content

Site Reliability Engineering

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/site-reliability-engineering

In short

Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.

What is site reliability engineering?

Site reliability engineering, or SRE, is an approach to running production systems that treats operations as a software problem. It was developed at Google in the early 2000s, when a team of software engineers was asked to run the company's services and chose to automate the work instead of doing it by hand. Site reliability engineers write code to deploy, scale, monitor, and repair systems, and they share responsibility for reliability with the developers who build the features.

SRE starts by measuring reliability from the user's point of view. Teams pick service level indicators (SLIs), such as the share of requests that succeed, set service level objectives (SLOs) for them, and treat the gap between the target and 100% as an error budget. While budget remains, teams ship quickly; when it runs out, they slow down releases and focus on stability. Other core practices include limiting toil, meaning repetitive manual work, sustainable on-call rotations, and blameless postmortems after incidents.

SRE is common at companies that run large online services, and many organizations have SRE or platform teams that support several product teams at once. A useful comparison is an airline's maintenance crew: they don't expect planes that never need repairs, but they define strict safety margins, follow checklists, study every incident, and build tools that catch problems before takeoff.

SRE is often confused with DevOps. DevOps is a broad culture and set of practices for bringing development and operations together, while SRE is one concrete way to put it into practice, with specific roles, metrics, and rules such as error budgets. A popular way to put it is class SRE implements DevOps: SRE is one specific implementation of the broader DevOps idea.

Key takeaways

  • SRE applies software engineering to operations and infrastructure work.
  • Reliability is measured with SLIs and SLOs that reflect what users experience.
  • An error budget balances the pace of new releases against stability.
  • SRE teams work to reduce toil, the repetitive manual work that can be automated.
  • Blameless postmortems turn incidents into lessons and follow-up fixes.

Readers ask

What is the difference between SRE and DevOps?

DevOps is a culture and set of practices that brings development and operations together. SRE is a specific way to put those ideas into practice, with dedicated engineers, reliability targets called SLOs, and error budgets that decide when to slow down releases.

What does a site reliability engineer do?

A site reliability engineer builds automation for deployment and scaling, sets up monitoring and alerts, joins an on-call rotation to respond to incidents, and leads postmortems. They also work with developers on designs so that new features are reliable and observable from the start.

What is toil in SRE?

Toil is manual, repetitive operational work that grows with the size of a service and has no lasting value, such as restarting stuck jobs by hand. SRE teams track toil and cap it, often at about half of their time, so the rest goes into engineering work that removes it.

Often compared

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings