Side by side
DevOpsvsSite Reliability Engineering
What is the difference between DevOps and SRE?
Updated 2 min read6 differences
In short
DevOps is a culture that unites development and operations to ship faster and safer; SRE is Google's practice of running production to reliability targets.
DevOps
Development and Operations
DevOps is a set of practices and a culture that brings software development and IT operations together to deliver software faster and more reliably.
Read the page on DevOpsSite Reliability Engineering
Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.
Read the page on Site Reliability EngineeringDevOps and Site Reliability Engineering compared
| Aspect | DevOps | Site Reliability Engineering |
|---|---|---|
| What it is | A culture and set of practices | A discipline and job role with defined methods |
| Origin | The DevOps movement of the late 2000s | Google, 2003 |
| Main goal | Deliver changes quickly and safely | Keep services reliable at an agreed level |
| How success is measured | Delivery metrics, such as deployment frequency and lead time | SLOs, error budgets and incident data |
| Typical work | CI/CD pipelines, automation, infrastructure as code | On-call, incident response, capacity planning, reducing toil |
| Relationship | The broad idea | One concrete way to put it into practice |
The difference, explained
DevOps grew in the late 2000s as a reaction to developers throwing code over a wall to a separate operations team. It is a culture and a set of practices, such as shared ownership, automation, continuous integration and delivery, and infrastructure as code, aimed at delivering changes often and reliably. Site Reliability Engineering, or SRE, began at Google in 2003, when Ben Treynor Sloss asked software engineers to run production systems.
The difference is breadth versus specifics. DevOps describes goals and habits without prescribing how to measure them. SRE brings concrete tools: service level objectives (SLOs) define how reliable a service must be, an error budget says how much unreliability is acceptable before releases slow down, toil is capped so that engineers have time to automate, and incidents end in blameless postmortems.
Google sums the relationship up as "class SRE implements interface DevOps": SRE is one way of practicing DevOps ideas. In many companies the roles overlap, with platform or DevOps teams building pipelines and tools, and SREs focusing on reliability, on-call duty, capacity and incident response for the most critical services.
A common misconception is that DevOps is a job title or a set of tools. Buying CI/CD tools or renaming the operations team doesn't change how development and operations work together, and SRE isn't operations with a new name either: it depends on engineers writing software to remove manual work.
Which one should you use?
Choose DevOps when…
- You want to change how development and operations work together.
- Your main problem is slow or risky releases.
- You are building pipelines and automation that many teams share.
Choose Site Reliability Engineering when…
- Your services are critical and reliability needs clear targets.
- You need a structured on-call and incident process.
- You want data, such as an error budget, to decide between new features and stability.
Readers ask
Is SRE a kind of DevOps?
In Google's words, SRE implements DevOps: it is one concrete, engineering-driven way of applying DevOps principles to running production.
Does a company need both?
Not necessarily as separate teams. Smaller companies often practice DevOps without dedicated SREs; larger ones add SRE teams for their most critical services.
What is an error budget?
The amount of unreliability an SLO allows. With a 99.9% availability target, about 43 minutes of downtime a month is the budget; once it is used up, the team works on stability before shipping new features.