Skip to content
Machine Behavior

On-call and escalation

Nothing pages anyone. Prometheus evaluates 17 alert rules, but the stack runs no Alertmanager and Prometheus has no alerting target, so a firing alert stays in Prometheus and Grafana until someone looks. One person answers for every service, and the agent sessions that build the platform are the responders. The one control that acts without a person is the deploy gate: a failed check stops the deploy and the previous build stays live.

Roles

RoleWhoWhat they do
OperatorStefan CoetzeeOwns every service in the catalog, decides, starts and stops agent sessions, holds the credentials
RespondersClaude Code sessions, one per track of workRead the records and dashboards, follow the runbooks, fix within their track, push through the gate, write the incident record
Gate ownerThe legislation-track sessionOwns the conformity checks; tier changes are the operator's decision
ControlThe conformity gateRuns before every deploy of the three sites; blocks on a failed check without anyone on call

A responder session works only while the operator has it running. There is no rotation, no second person and no response time agreed with anyone.

How a problem is noticed today

  1. The gate fails. The deploy does not run, the run is red on GitHub Actions, the status page shows the site as degraded and the Website deploys dashboard shows the failing check. Readers keep the previous build.
  2. An alert fires. It shows in Prometheus and on the internal Grafana dashboards. Nobody is told.
  3. A session or the operator sees it while working: a dashboard, a failed search, a wrong number.
  4. A reader reports it as a GitHub issue; it reaches the board after triage.

Of the four incidents on the status page, the first route found one (the blocked deploy of 2026-10-08) and the third route found the other three.

The alert rules

GroupRulesSeverity
Local LLM SLOsLocalLLMAvailabilityBudgetBurnFast, LocalLLMLatencyBudgetBurnFast, LocalLLMAvailabilityBudgetBurnSlowpage, page, ticket
Local LLM symptomsOllamaDown, OllamaExporterDown, ModelSpilledToCPUpage, page, info
Claude Code spendClaudeCodeSpendSpiketicket
SearXNGSearXNGDown, SearXNGProxyEgressDown, SearXNGTunnelDown, SearXNGEngineFailing, SearXNGDegradedSearches, SearXNGSlowSearchespage, then five ticket
AWS costAWSCostForecastOverBudget, AWSCostMonthToDateOver80, AWSCostDailySpike, AWSCostExporterStaleticket

Five rules carry severity: page, eleven ticket and one info. Outside the stack, the AWS monthly budget in AWS Budgets sends an e-mail at 80% of actual spend; it is the one cost alert that reaches a person (Budgets and alerts). Conditions and tests: Alerts and SLOs.

Response

  1. Open a record in incidents/ with an investigating update and push; the status page shows it after the deploy. See Status page.
  2. Find the service in the catalog and follow its runbook.
  3. Fix through the gate. A fix that cannot pass the gate does not go out.
  4. Move the record through identified, monitoring and resolved.
  5. For anything readers saw, write the cause and the fix into the docs and link it from the record.

Order of work follows the service tier: tier 1 first, tier 2 the same day, tier 3 in the next working session.

Escalation

FromToWhen
Responder sessionThe operator, in the sessionAnything outside the session's track, anything that needs a credential, a cost, a change to the gate or a public statement
OperatorThe gate ownerA false positive or a rule change in the gate
OperatorThe providerGitHub for Pages and Actions, Anthropic for the API, the upstream search engines for blocks

Proposed routing

This is a proposal and is not wired. Sending a page to a phone or an address needs the operator's go and a receiver the operator chooses.

# alertmanager.yml (proposal)
route:
  receiver: board            # default: a ticket on the board
  group_by: [alertname]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 24h
  routes:
    - matchers: [severity="page"]
      receiver: operator-push
      repeat_interval: 1h
      active_time_intervals: [waking-hours]
    - matchers: [severity="page"]   # outside waking hours a page becomes a ticket
      receiver: board
    - matchers: [severity="info"]
      receiver: blackhole
    - matchers: [alertname="Watchdog"]
      receiver: heartbeat
      repeat_interval: 5m
inhibit_rules:
  - source_matchers: [alertname="OllamaDown"]
    target_matchers: [alertname=~"LocalLLM.+|ModelSpilledToCPU|OllamaExporterDown"]
  - source_matchers: [alertname="SearXNGDown"]
    target_matchers: [alertname=~"SearXNG.+"]
time_intervals:
  - name: waking-hours
    time_intervals:
      - times: [{start_time: "08:00", end_time: "22:00"}]
        location: Europe/Berlin
receivers:
  - name: operator-push      # a push service on the operator's phone; not chosen
  - name: board              # a webhook that opens or updates a GitHub issue with type and area labels
  - name: heartbeat          # an external check that alarms when the Watchdog stops arriving
  - name: blackhole

What it would change:

  • The five page rules reach a person during waking hours; at night they wait as tickets, which matches one operator.
  • Tickets land on the board as issues, so an alert has an owner and a history next to the other work.
  • An always-firing Watchdog rule and an outside heartbeat check would show when the alerting path itself is down. Without it, a dead Prometheus and a quiet night look the same.
  • Inhibition keeps one cause from raising five alerts: a down Ollama silences the SLO burns behind it.

Wiring it needs an Alertmanager service in the compose file, an alerting block in prometheus.yml, a Watchdog rule, promtool tests for the routes, and the operator's choice of receivers.

Built from scripts/docs by build_docs.py.