The principle
Watch the four golden signals of each user-facing service: latency, traffic, errors and saturation. Alert on symptoms, what users get, and keep causes for dashboards and tickets, because one cause can show as many symptoms and many causes never reach a user. Every page must be urgent, actionable and new to the person who gets it. Black-box checks from outside show what users see now; white-box metrics from inside show what is about to break.
Source: Symptoms over Causes for Alerting under Observability; Google's Monitoring Distributed Systems and Alerting on SLOs.
On this platform
- Symptom alerts where an SLI exists. The three Local LLM burn-rate alerts read the error ratio and the slow-first-token ratio that callers get.
SearXNGDegradedSearchesfires when the search script returned a degraded result 3 or more times in 15 minutes, andSearXNGSlowSearcheswhen search latency p95 stays above 8 s (Alerts and SLOs). - Most cause alerts are tickets. Tunnel, proxy, engine and cost-exporter failures are tickets. Three cause alerts carry severity page:
OllamaDown,OllamaExporterDownandSearXNGDown. The proposed routing would silence the SLO burns behind a down Ollama, so one cause raises one alert (On-call and escalation). - White-box telemetry. The observability stack runs Prometheus, Loki, Tempo and Grafana, fed by exporters for the model server and the deploy pipeline (Architecture), the search instance (Alerts and SLOs) and AWS cost (Budgets and alerts). Five dashboards are public (Public dashboards), among them Website deploys.
- Absence is not zero. The SearXNG exporter pushes from a laptop, so a sleeping laptop leaves a gap in the series. The SearXNG rules read pushed values or failure ratios with a minimum number of calls, and none reads
uporabsent(). - A meter that was wrong. Claude Code's cost counters reset each time a parallel session exported a lower total, and Prometheus summed the resets into USD 1.97M for a week of USD 1,003. The spend alert was in firing state 73% of the time on a false signal. Spend is now summed from per-request events (incident record, Counter resets from parallel sessions).
- The reader's view. The status page shows each service from the published records, with open and past incidents.
State
Partial. Two services have symptom alerts and the telemetry is wide. Three parts are missing. No alert is delivered to anyone (#36). No black-box probe checks the three sites from outside, so the most visible services have no symptom signal of their own (#37). And the incident records do not say how each incident was detected; apart from the deploy the gate blocked, all were found by someone working at the time (#41).
Services that name it
In a company of 10 to 200 people
- Give each user-facing service the four golden signals on one dashboard, and one black-box probe from outside your own network.
- Page on symptoms tied to an SLO; send causes to tickets and dashboards.
- Review every page after the fact: was it urgent, was it actionable, was it new? Delete or demote the rules that fail.
- Test the alert rules like code, and alert on the alerting path itself.