Runbooks

These runbooks cover the traps found while running the stack. Each one gives the symptom, the cause, a check and the fix. Commands run on the home server from the repository directory unless a step says otherwise.

RunbookSymptom
Backfilled runs missing in LokiThe Website deploys dashboard shows no history after the deploy-exporter starts with an empty state
Stat panel shows No dataA stat panel is empty over a long range while the metric exists
Counter resets from parallel sessionsClaude Code spend computed from Prometheus counters reads in the millions of USD
Share a dashboard publiclyTask: publish a dashboard at grafana.scoetzee.de

First checks

These commands show whether each service runs and whether the exporters serve metrics:

docker compose ps
docker compose logs --tail 50 deploy-exporter
curl -s http://localhost:9201/metrics | grep site_deploy_exporter
curl -s http://localhost:11435/metrics | grep ollama_up

Health signals

SignalQueryHealthy value
Deploy-exporter poll agetime() - site_deploy_exporter_last_poll_timestamp_secondsClose to the 60 s poll interval
Deploy-exporter poll errorssite_deploy_exporter_errors_totalFlat; it counts failed repository polls since the exporter started
Ollama reachableollama_up1; the OllamaDown alert fires after 2 minutes at 0
Proxy scrapedup{job="ollama-exporter"}1; the OllamaExporterDown alert fires after 2 minutes at 0
Claude Code spendsum(model:claude_code_cost_usd:sum1h)Under USD 20 per hour, the ClaudeCodeSpendSpike threshold

The deploy-exporter writes one log line per run it reads, with site, run number, outcome, duration and gate result. Failure lines carry poll failed, log fetch failed, loki flush failed, probe fetch failed or state save failed, followed by the site where one applies and the error.

Alert rules and thresholds: Alerts and SLOs.