Skip to content
Machine Behavior

SRE is a role inside Ops

The principle

SRE is a specialisation within Ops, the outer ring of a concentric model: system and network admins at the core, then cloud engineers, platform engineers and SRE, each ring holding the skills of the rings inside it. DevOps, FinOps, SecOps and DataOps are overlaps between Ops and another domain. An SRE team cut off from Ops ends up as a second platform team or as developers who no longer run anything.

Source: the manifesto, with the team-design side in The Disappearing Full-Stack Ops Engineer. Google's introduction starts from the other side, with software engineers asked to design an operations function. In the manifesto, Ops comes first and the engineering is added to it.

On this platform

  • One operator, every ring. Stefan Coetzee owns all twelve services in the catalog. Each ring has work on the platform: the home server and its network (core), the small cloud proxy (cloud), the gate, the deploy job and the generators (platform), the SLOs, alert rules and incident records (SRE).
  • Owner and operator are separate fields. A service can name who runs it day to day apart from who answers for it. The conformity gate names an operator, the legislation-track agent session. The agent sessions are the responders, one per track of work, under the operator (On-call and escalation).
  • The overlaps run as practice. FinOps has its own space with a cost model, showback and budgets. SecOps shows in the access model, the rule against third-party requests on page load (ADR-0010) and the redaction step of the vault import (Docs tree). DevOps is the gate inside the deploy pipeline.
  • Tiers order the work. Tier 1 first, tier 2 the same day, tier 3 in the next working session (Service catalog).

State

Partial. The roles are named and every ring is covered, by one person. No second person can take a ring, a page or a review, so the model is visible in the work and a team does not exist yet. Two security items are blocked: security headers on the hosting decision (#26) and security.txt on a contact route for vulnerability reports (#12).

Services that name it

In a company of 10 to 200 people

  • Map the people you have to the rings before you hire. A company of 30 often has developers doing platform work and nobody at the core ring; that ring is missed in the first outage that needs it.
  • Hire SRE into Ops, with standby in the job description, and pair each SRE with the developers of the services they run.
  • Give every service an owner and, where it differs, an operator, in a catalog the build checks.
  • Name FinOps and SecOps as duties of named people early. They become teams later.

Open work

Built from scripts/docs by build_docs.py.