Nothing pages anyone. Prometheus evaluates 17 alert rules, but the stack runs no Alertmanager and Prometheus has no alerting target, so a firing alert stays in Prometheus and Grafana until someone looks. One person answers for every service, and the agent sessions that build the platform are the responders. The one control that acts without a person is the deploy gate: a failed check stops the deploy and the previous build stays live.
Roles
| Role | Who | What they do |
|---|---|---|
| Operator | Stefan Coetzee | Owns every service in the catalog, decides, starts and stops agent sessions, holds the credentials |
| Responders | Claude Code sessions, one per track of work | Read the records and dashboards, follow the runbooks, fix within their track, push through the gate, write the incident record |
| Gate owner | The legislation-track session | Owns the conformity checks; tier changes are the operator's decision |
| Control | The conformity gate | Runs before every deploy of the three sites; blocks on a failed check without anyone on call |
A responder session works only while the operator has it running. There is no rotation, no second person and no response time agreed with anyone.
How a problem is noticed today
- The gate fails. The deploy does not run, the run is red on GitHub Actions, the status page shows the site as degraded and the Website deploys dashboard shows the failing check. Readers keep the previous build.
- An alert fires. It shows in Prometheus and on the internal Grafana dashboards. Nobody is told.
- A session or the operator sees it while working: a dashboard, a failed search, a wrong number.
- A reader reports it as a GitHub issue; it reaches the board after triage.
Of the four incidents on the status page, the first route found one (the blocked deploy of 2026-10-08) and the third route found the other three.
The alert rules
| Group | Rules | Severity |
|---|---|---|
| Local LLM SLOs | LocalLLMAvailabilityBudgetBurnFast, LocalLLMLatencyBudgetBurnFast, LocalLLMAvailabilityBudgetBurnSlow | page, page, ticket |
| Local LLM symptoms | OllamaDown, OllamaExporterDown, ModelSpilledToCPU | page, page, info |
| Claude Code spend | ClaudeCodeSpendSpike | ticket |
| SearXNG | SearXNGDown, SearXNGProxyEgressDown, SearXNGTunnelDown, SearXNGEngineFailing, SearXNGDegradedSearches, SearXNGSlowSearches | page, then five ticket |
| AWS cost | AWSCostForecastOverBudget, AWSCostMonthToDateOver80, AWSCostDailySpike, AWSCostExporterStale | ticket |
Five rules carry severity: page, eleven ticket and one info. Outside the stack, the AWS monthly budget in AWS Budgets sends an e-mail at 80% of actual spend; it is the one cost alert that reaches a person (Budgets and alerts). Conditions and tests: Alerts and SLOs.
Response
- Open a record in
incidents/with aninvestigatingupdate and push; the status page shows it after the deploy. See Status page. - Find the service in the catalog and follow its runbook.
- Fix through the gate. A fix that cannot pass the gate does not go out.
- Move the record through identified, monitoring and resolved.
- For anything readers saw, write the cause and the fix into the docs and link it from the record.
Order of work follows the service tier: tier 1 first, tier 2 the same day, tier 3 in the next working session.
Escalation
| From | To | When |
|---|---|---|
| Responder session | The operator, in the session | Anything outside the session's track, anything that needs a credential, a cost, a change to the gate or a public statement |
| Operator | The gate owner | A false positive or a rule change in the gate |
| Operator | The provider | GitHub for Pages and Actions, Anthropic for the API, the upstream search engines for blocks |
Proposed routing
This is a proposal and is not wired. Sending a page to a phone or an address needs the operator's go and a receiver the operator chooses.
# alertmanager.yml (proposal)
route:
receiver: board # default: a ticket on the board
group_by: [alertname]
group_wait: 30s
group_interval: 5m
repeat_interval: 24h
routes:
- matchers: [severity="page"]
receiver: operator-push
repeat_interval: 1h
active_time_intervals: [waking-hours]
- matchers: [severity="page"] # outside waking hours a page becomes a ticket
receiver: board
- matchers: [severity="info"]
receiver: blackhole
- matchers: [alertname="Watchdog"]
receiver: heartbeat
repeat_interval: 5m
inhibit_rules:
- source_matchers: [alertname="OllamaDown"]
target_matchers: [alertname=~"LocalLLM.+|ModelSpilledToCPU|OllamaExporterDown"]
- source_matchers: [alertname="SearXNGDown"]
target_matchers: [alertname=~"SearXNG.+"]
time_intervals:
- name: waking-hours
time_intervals:
- times: [{start_time: "08:00", end_time: "22:00"}]
location: Europe/Berlin
receivers:
- name: operator-push # a push service on the operator's phone; not chosen
- name: board # a webhook that opens or updates a GitHub issue with type and area labels
- name: heartbeat # an external check that alarms when the Watchdog stops arriving
- name: blackhole
What it would change:
- The five page rules reach a person during waking hours; at night they wait as tickets, which matches one operator.
- Tickets land on the board as issues, so an alert has an owner and a history next to the other work.
- An always-firing
Watchdogrule and an outside heartbeat check would show when the alerting path itself is down. Without it, a dead Prometheus and a quiet night look the same. - Inhibition keeps one cause from raising five alerts: a down Ollama silences the SLO burns behind it.
Wiring it needs an Alertmanager service in the compose file, an alerting block in prometheus.yml, a Watchdog rule, promtool tests for the routes, and the operator's choice of receivers.
Related
- Alerts and SLOs: every rule and its condition.
- Runbooks and the Engineering runbooks.
- Access model: who can reach what.