Skip to content
Machine BehaviorCTO
Menu

Every incident ends in a record

The principle

An incident ends when the service works again and the record is written. A blameless postmortem assumes everyone acted reasonably on the information they had, and asks what in the system let the failure happen and what would catch it earlier. The record holds the impact, the timeline, the cause and follow-up work with owners, and the lesson turns into a runbook, a test or a fix.

Source: Incident Management and Blameless Postmortem Discipline; Google's Postmortem Culture: Learning from Failure.

On this platform

  • Records with stages. Each incident is a YAML file in incidents/ with impact, services, summary and dated updates through investigating, identified, monitoring and resolved. The build stops when the stages go backwards or a service is not in the catalog, and the status page renders open and past incidents (Status page).
  • Write-ups linked from the record. The phantom spend record links a FinOps case study, a runbook, the budgets page and an open follow-up ticket (#29). The tag pages record names the guard the crawler kept afterwards. The SearXNG record, still open, lists every change made and the choice that remains.
  • A control working is also a record. The blocked deploy of 2026-10-08 had no reader impact; it is kept because a blocked deploy is the case the gate exists for, and its runbook was written the same day.
  • Runbooks from failures. Each Engineering runbook covers one failure that has happened, with how to see it, the cause and the fix (Runbooks); the Observability runbooks follow the same rule (Observability runbooks).
  • Records name mechanisms. The records say what broke and why; none names a person at fault. The handbook's own incident records from earlier work follow the same form and name their sources.

State

Partial. Every incident on the platform has a record and a write-up; three are resolved, and the SearXNG record is still open at monitoring. Three parts are missing. The records do not say when and how each incident was detected; three of four do not record the time of the first report or check (#41). The incident records in the handbook have no index and no criteria for when a postmortem is written (#22), and there is no template for a write-up (#20). Follow-up work has a ticket for one of the four incidents.

Services that name it

In a company of 10 to 200 people

  • Write a record for every incident that reached a customer and for every near miss a control caught, within five working days.
  • Use one template: impact, timeline with detection time, contributing causes, what went well, follow-up items with owners and dates.
  • Review the records in a monthly meeting open to the whole engineering team, and track follow-up items on the same board as feature work.
  • Ban names in the cause section. Ask what made the action reasonable at the time.

Open work

Built from scripts/docs by build_docs.py.