Skip to content
Machine Behavior

Someone carries the pager

The principle

If one test decides whether someone does SRE work, it is standby. SRE carries the pager and is the first rotation for production outages; platform and ops teams stand behind it on their own rotations. A team with the title and no standby is a platform team. Google's book adds the limits that keep standby sustainable: a cap on incidents per shift, time to follow up each one, and most of an SRE's time kept for engineering.

Source: the standby section of the manifesto; Google's Being On-Call.

On this platform

  • Rules exist, delivery does not. Prometheus evaluates 17 alert rules: 5 with severity page, 11 ticket and 1 info (Alerts and SLOs). The stack runs no Alertmanager, so a firing alert stays in Prometheus and Grafana until someone looks (On-call and escalation).
  • One path reaches a person. The AWS monthly budget in AWS Budgets sends an e-mail at 80% of actual spend. It sits outside the stack (Budgets and alerts).
  • One control acts at any hour. The conformity gate stops a failing deploy and keeps the previous build live without anyone on call. On 2026-10-08 it blocked a deploy, and the fix passed 81 seconds later (incident record).
  • How incidents were found. Of the four records on the status page, the gate found one, the blocked deploy. Someone who was working at the time found the other three, and none of those three records the time of the first report or check.
  • A routing proposal. The on-call page holds a routing: pages to the operator's phone in waking hours, tickets onto the board, an always-firing Watchdog with an outside heartbeat, and inhibition so one cause raises one alert.

State

Gap. Nobody is on standby and no alert rule reaches anyone. The gate covers the largest tier-1 risk, a bad page in front of readers, without a person. Every other failure stays unseen until someone looks: a down search instance, a model server that stopped, a spend spike at night. The proposal is wired once the operator chooses the receivers (#36).

Services that name it

No service in the catalog names this principle.

In a company of 10 to 200 people

  • Start with one rotation for the services customers meet. Each person should be on call at most one week in four; with fewer people, share one rotation across teams.
  • Page only on symptoms users feel, with a runbook link in every page. Everything else becomes a ticket with an owner.
  • Pay for standby or give the time back, and review the pager load every month.
  • Add a dead man's switch: an always-firing alert that an outside check expects, so a broken alert path raises an alarm too.

Open work

Built from scripts/docs by build_docs.py.