Home Docs SRE Handbook Principles in practice
Principles in practice
owner Stefan Coetzee created 2026-10-09 updated 2026-10-09 reviewed 2026-10-09 review by 2027-01-07 3 min read explanation
This section ties the SRE framework of this handbook to the platform it runs on. Each page states one principle and its source, shows how the platform behind machinebehavior.io applies it with links to the live records, gives the state of that practice, and says what a company of 10 to 200 people would do with it. Each gap has an issue on the board .
Stefan Coetzee's framework comes first. His manifesto defines SRE as a role inside Ops, with standby as its test, and puts the substrate principle under the ten pillars . Part II of Google's Site Reliability Engineering book is the reference the framework extends: embracing risk, service level objectives, eliminating toil, monitoring distributed systems, automation, release engineering and simplicity.
The principles
Principle Source State Services that name it Truth, verified, working Manifesto , the substrate principlepartial Conformity gate , Docs build , machinebehavior.io SRE is a role inside Ops Manifesto , the concentric modelpartial Agent sessions , Conformity gate Someone carries the pager Manifesto , the standby differentiator; Google SRE book ch. 11 gap none Operations work done as software Manifesto , SRE = Ops + software engineering; Google SRE book ch. 7 in place Conformity gate , Deploy exporter , Docs build , Map crawler , Observability stack Risk is a budget Pillar 1, error budgets ; Google SRE book ch. 3 partial Conformity gate , Local LLM Service level objectives Pillar 1, reliability ; Google SRE book ch. 4 gap Local LLM Monitor symptoms, page on user pain Pillar 3, symptoms over causes ; Google SRE book ch. 6 partial Deploy exporter , Local LLM , Observability stack , SearXNG Eliminate toil Pillar 10, toil reduction ; Google SRE book ch. 5 partial Deploy job , Docs build , Map crawler Every change passes the same gate Pillar 6, CI/CD and deployment ; Google SRE book ch. 8 partial Conformity gate , Deploy exporter , Deploy job , machinebehavior.io , tychat.io , uncovertechtalent.com Simplicity Google SRE book ch. 9 ; Idempotence in pillar 5in place Deploy job , machinebehavior.io , tychat.io , uncovertechtalent.com If it is not in git, it does not exist Pillar 5, infrastructure as code ; Google SRE book ch. 8 , configuration managementpartial Docs build , machinebehavior.io , Observability stack Every incident ends in a record Pillar 4, blameless postmortems ; Google SRE book ch. 15 partial Agent sessions , SearXNG Cost per unit of work Pillar 9, unit economics ; FinOps Foundation framework partial Agent sessions , Observability stack Agents in production are ops work Manifesto , AI-era extensionspartial Agent sessions
How to read the state
State Meaning in place The platform applies the principle, and a reader can check it on a live page or record partial Some of it runs; the missing part is named on the page and on the board gap The platform does not apply it yet; the page says why and links the work
The state is set by hand at each review, against the live pages. The build stops when the state in a page's front matter and the state in its text disagree. Each service page lists the principles it applies, under "SRE principles applied", from the principles: field of its YAML file ; the table under The principles reads the same field.
The platform these pages describe
One operator, Stefan Coetzee, and Claude Code sessions that write most of the code and the pages. Three static sites on GitHub Pages behind one conformity gate ; an observability stack and a local LLM on a home server; a search instance on a laptop; a small cloud proxy. Twelve services in the catalog , four incident records on the status page , the cost side in the FinOps space . At this size some principles cost little and some do not pay yet; each page says which.
The five-minute founder tour walks the same evidence in six stops.
Child pages Truth, verified, working An operational claim must be true, backed by a probe in its time window, and about a system that works. The gate, the status page and the review labels apply it here; numbers written into the docs still go stale without a probe. SRE is a role inside Ops SRE is Ops with software engineering applied, the outer ring of a concentric model, with FinOps, SecOps and DevOps as overlaps. Here one operator holds every ring, agent sessions do the work, and the catalog names an owner for each service. Someone carries the pager The test for SRE work is standby, a person who is told when production breaks and acts. Here 17 alert rules reach nobody, the gate is the one control that acts at any hour, and three of four incidents were found by someone working at the time. Operations work done as software SRE is operational knowledge plus software engineering, with automation in place of manual steps and tools in place of one-off scripts. Here generators build the catalog, status page, docs, map, board and changelog from sources in git, and each build checks its input. Risk is a budget A target of 100% reliability costs more than users can notice; the gap between the target and 100% is an error budget for change. Here tiers set the order of response and the Local LLM has burn-rate alerts, but no policy says what happens when a budget is spent. Service level objectives An SLO is a target for a measured indicator of what users get, and the basis for alerts and error budgets. Here one of twelve services has SLOs, the Local LLM; the five tier-1 services, which readers meet or which decide whether a deploy goes out, have none. Monitor symptoms, page on user pain Monitoring should say what is broken for users before it says why, and a page should reach a person only when they must act. Here burn-rate and degraded-search alerts read what users get, five dashboards are public, and no alert is delivered. Eliminate toil Toil is manual, repetitive, automatable work that grows with load and leaves nothing behind; SRE caps it so engineering time remains. Here generators and the deploy job removed most of it; a rebase before every push, hand deploys of the observability stack and the screenshot retakes remain, and nobody measures the share. Every change passes the same gate Release engineering makes every change go through one repeatable, recorded path, with builds that give the same output wherever they run. Here one workflow gates and deploys three sites and records every run; the generated pages are built on the author's machine, and the observability stack has no pipeline. Simplicity Every line of code and every moving part is a liability, so prefer boring technology, small interfaces and deleted code. Here three static sites, standard-library Python, no third-party requests on load and one navigation source keep the platform small enough for one operator. If it is not in git, it does not exist Infrastructure and configuration are declared in version control and applied the same way every time, so the repository describes the running system. Here the sites, the catalog, the incidents, the gate rules, the alert rules and the dashboards are in git; the hosts, the dashboard sharing and the GitHub settings are not. Every incident ends in a record Each incident is written up without blame, with impact, timeline, cause and follow-up work, and the lesson becomes a runbook or a fix. Here four incident records sit on the status page with stages, write-ups and tickets; detection times, a template and an index are missing. Cost per unit of work Cost is an operational signal read per unit of output, metered like latency and alerted like errors. Here a FinOps space meters each component, prices spend per call, commit and deploy, and sets alert thresholds from data; only the AWS budget e-mail reaches a person. Agents in production are ops work An AI agent in production is a service and needs what services need: owners, SLOs, paging, runbooks, cost control and canaries. Here the agent sessions are a catalogued service that pushes through the same gate, with metered spend and an incident record; they have no SLOs and no canary set.
principles sre practice
Linked from
Built from scripts/docs by build_docs.py .