Observability
Prometheus, Loki and Grafana behind the public dashboards, the alert rules and SLOs, and the status page that reads the published records.
6 Inside pages, 21 docs pages, 1 post, 12 tickets
Inside 6
- Mission Control
The live operations view of the platform behind machinebehavior.io on one screen: gate and deploy state of the three sites, the deploy feed, open incidents, the conformity gate result and the public Grafana dashboards.
- Inside
One front page for Stefan Coetzee's published work: the latest pieces across machinebehavior.io, tychat.io, uncovertechtalent.com and Substack, live deploy and gate status of the three sites, experiment results, the live Grafana dashboards and the tools.
- Status
Status of the 12 services behind machinebehavior.io, read from the gate records and the deploy feed, and the incident history with the stages Investigating, Identified, Monitoring and Resolved.
- Observability stack
One Docker Compose project on a home server: Alloy, Prometheus, Loki, Tempo and Grafana, with 17 alert rules and the public dashboards behind a reverse proxy that passes only the shared paths. No Alertmanager runs, so no alert pages anyone.
- Deploy exporter
A standard-library Python service that reads the GitHub Actions runs of the three sites, writes each run, step and gate check to Loki and serves the latest state as Prometheus metrics. Feeds the Website deploys dashboard.
- SearXNG
The self-hosted metasearch instance every research agent searches through, with per-engine pacing and suspensions after rate limits. An exporter pushes engine health to the observability stack; six alert rules watch it.
Docs 21
Engineering 1
- ADR-0026: Mission Control, the live operations view inside the portalOne screen at /inside/mission-control/ shows the gate and deploy state of the three sites, the deploy feed, open incidents, this site's gate result and the public Grafana dashboards as tabs that load on a click, with a slot for the internal AWS spend panel.
Observability 14
- ObservabilityMetrics, logs and traces for Claude Code, a self-hosted Ollama server, its host and three website deploys, collected by one Docker Compose stack on a home server.
- Alerts and SLOsThe recording rules, service level objectives and 17 alerts that Prometheus and the Loki ruler evaluate, with thresholds from the rule files.
- ArchitectureHow the collectors, the three stores, Grafana and the public front door of the observability stack connect, and how the stack is deployed.
- Backfilled runs missing in LokiWhat to do when the Website deploys dashboard shows no older runs after the deploy-exporter backfills from GitHub Actions.
- Counter resets from parallel sessionsWhy Claude Code's cost counters in Prometheus report spend in the millions of USD, and how the stack reads spend and tokens from per-request events in Loki.
- Dashboards as codeHow grafana/build.py generates every dashboard as JSON, how the public cuts derive from the internal dashboards, and how a change reaches Grafana.
- Deploy exporterA standard-library Python service that reads the GitHub Actions deploy runs of three websites, writes them to Loki and serves the latest state as Prometheus metrics on port 9201.
- Loki streamsThe Loki streams the stack writes, with their labels and line fields, for website deploy records and for Claude Code events.
- Metrics referenceEvery Prometheus metric the deploy-exporter and the ollama-exporter serve, with type, labels and meaning, taken from the exporter source.
- On-call and escalationThe real on-call model: one operator, agent sessions as responders, the deploy gate as the one control that acts on its own, and 17 alert rules that page nobody because no Alertmanager runs. With a proposed routing.
All 14 pages in Observability
- Public dashboardsThe five Grafana dashboards shared at grafana.scoetzee.de, what their panels show, and how each public cut differs from the internal dashboard.
- RunbooksStep-by-step fixes for the known traps in the observability stack: missing backfill in Loki, empty stat panels, Claude Code counter resets and sharing a dashboard.
- Share a dashboard publiclyHow to build a variable-free public cut of a dashboard, provision it, and turn on Grafana public sharing so it loads at grafana.scoetzee.de.
- Stat panel shows No dataWhy Grafana stat panels on young series show No data over long time ranges, and how build.py switches them to instant queries.
FinOps 1
- Budgets and alertsThe cost controls that exist (the Claude Code spend alert, the AWS monthly budget with its internal dashboard and four alert rules, the eval harness caps), why only the AWS Budgets e-mail reaches a person, and what a budget alert for Claude Code would look like.
SRE Handbook 5
- CPU Cache Hierarchy and Speculative ExecutionMain memory is the source of truth.
- Error Budgets as Reliability CurrencyAn error budget is the only mechanism that makes the reliability-vs-velocity trade-off concrete.
- Monitor symptoms, page on user painMonitoring should say what is broken for users before it says why, and a page should reach a person only when they must act. Here burn-rate and degraded-search alerts read what users get, five dashboards are public, and no alert is delivered.
- Service level objectivesAn SLO is a target for a measured indicator of what users get, and the basis for alerts and error budgets. Here one of twelve services has SLOs, the Local LLM; the five tier-1 services, which readers meet or which decide whether a deploy goes out, have none.
- Symptoms over Causes for AlertingPage when users are affected.
Posts 1
- Success Is the Engine RunningThe modern world was built on explosions: contained, timed and measured ones. That is what an engine is. Language model output is the fire; the harness, the hook and the far-end gauge are the engine.
Tickets 12
Issues on the board, open first, as of the last build.
- #41 Incident records: when and how each incident was detected
- #38 Error budget policy: what happens when a service spends its budget
- #37 SLOs for the tier-1 services and the agent sessions
- #36 Alert routing: Alertmanager, a Watchdog and receivers chosen by the operator
- #29 Cost counters: drop them or alert on their resets (incident follow-up)
- #25 Retention statement for the metrics and log stores
- #22 Postmortem index
- #42 Mission Control at /inside/mission-control/
- #18 On-call and escalation page, and Alertmanager
- #15 Status page at /inside/status/
- #6 SearXNG search engines dashboard on Inside
- #5 Remove internal host details from the public dashboards
Other topics
Built by build_hubs.py from site/topics.yml, the docs labels, the tag pages on the map and the board. Machine-readable: topics.json.