Skip to content
Machine BehaviorCTO
Menu

Service level objectives

The principle

Pick a few indicators that describe what users get from a service (the share of requests that succeed, the share answered fast enough), measure them as ratios of good events to all events, and set a target for each over a window. The target is the SLO; an SLA adds a consequence for missing it and is set looser than the SLO. Without SLOs there is no error budget, and alerts fall back on causes in place of user pain. Scalability and performance are the same discipline in the capacity and latency dimensions.

Source: Reliability, the first of the ten pillars; Google's Service Level Objectives and Implementing SLOs.

On this platform

  • The worked example. The Local LLM has two SLOs with their good events written out: availability 99% (a request without an upstream error) and latency 95% (a streamed chat request with a first token within 4 s). Recording rules compute both ratios over 5 minutes, 30 minutes, 1 hour and 6 hours, and burn-rate alerts read them (Alerts and SLOs). The public dashboard shows the requests, errors and time to first token behind them.
  • The rest of the catalog. The other eleven service pages say "None defined" under Service level objectives, and the status page reads their state from published records: the gate record and the deploy feed for the sites, the gate and the deploy job, the snapshot time for the map crawler. The other six show "no live check from this page", and only an open incident changes their state.
  • Data that exists for SLIs. The deploy exporter records every Actions run with its steps and gate checks, enough for a deploy SLO. Per-request Claude Code events in Loki carry model, tokens and cost, enough for an agent spend SLO (Cost model). The SearXNG exporter records search latency and degraded results (incident record).

State

Gap. One of twelve services has SLOs, and it is a tier-3 research tool. The three sites, the gate and the deploy job have none, so nobody can say how reliable they were last month. The sites also lack the probe an availability SLI needs: the map crawl in the deploy job fetches every page on each deploy, but nothing requests them every minute and records the answers. The issue lists candidate indicators to check against the data before any target is set (#37).

Services that name it

In a company of 10 to 200 people

  • Start with the one or two user journeys that pay the bills, one availability and one latency SLI each, measured as close to the user as you can.
  • Set the first target from a month of data, a little below what the system already does, and tighten it only when users ask.
  • Put the SLO in the service catalog next to the owner, so every service page answers "how reliable is it supposed to be?".
  • Review the SLOs each quarter with the product owner; drop an SLO nobody acts on.

Open work

Built from scripts/docs by build_docs.py.