Skip to content
Machine Behavior

Status page

The status page shows the current state of every service in the catalog and the incident history. It is built by scripts/build_inside.py from services/*.yml and incidents/*.yml; the service states are read in the visitor's browser.

Where the state comes from

The page reads only records the site already publishes, so a visit sends no request to a third party.

RecordUsed forWritten by
https://<site>/conformity/latest.json of each siteGate pass or fail, time of the last runThe conformity job of each site
/inside/deploys.jsonResult and push-to-live time of the last workflow run per public repositoryThe deploy job of this site
map_generated in /inside/search.jsonAge of the map snapshotThe deploy job, after the map refresh
An open incident with reader impactMarks the services it names as degradedincidents/*.yml

The status block of a service file names its records (gate, deploys). A service without one shows "No live check" and links its dashboard. The observability stack, SearXNG, the local LLM and the agent sessions have no record on this site; Grafana and the alert rules watch them. See Alerts and SLOs.

StateRule
OperationalThe gate record passes and the last run on the deploy feed succeeded
DegradedThe gate fails (new deploys blocked, the previous build stays live), the last run failed, the map snapshot is older than eight days, or an open incident with reader impact names the service
OutageA site does not answer with its gate record
No live checkNo record for the service on this site

Incidents

One YAML file per incident in incidents/, named YYYY-MM-DD-short-name.yml.

FieldRequiredValues
idyesSame as the file name
title, summaryyesPlain text, blameless: what happened to the system, never who
servicesyesService ids from the catalog
impactyesnone, minor, major, critical; none does not mark services as degraded
startedyesYYYY-MM-DD or YYYY-MM-DDTHH:MMZ (UTC)
resolvedwhen resolvedSame format; set exactly when the last update is resolved
updatesyesA list of stage, at, text; stages only go forward: investigating, identified, monitoring, resolved
follow_upnoWhat changed so it does not happen again
postmortemnodoc:space/slug links to the write-up and runbooks
ticketsnoGitHub issue numbers for the owned follow-up work, for example [29]; each links to its card on the board

A time that was not recorded is written as a date alone and shown as "time not recorded". Only times with a record behind them (a commit, a run, a log line) carry a clock time.

Open, update, close

  1. Open: add the file with one investigating update, run python3 scripts/build_inside.py, run the gate dry run and push.
  2. Update: append an update with the next stage, or another update in the same stage, and push.
  3. Close: append a resolved update, set resolved, link the write-up in postmortem, and push.

The status page is as fresh as the last deploy. The incident text is in the page source, so the gate scans it like any other page.

Rules for the text

  • No private addresses, host names, personal details or employer names.
  • Name the mechanism and the fix; no person is the cause.
  • Numbers come from a record: a run, a commit, a query result.

Built from scripts/docs by build_docs.py.