4. Incident Management

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-04-25 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

"It's not about preventing all failures. It's about recovering fast."

What is Incident Management?

The structured approach to identifying, responding to, resolving, and learning from service disruptions.

Incident Lifecycle

Detection โ†’ Triage โ†’ Response โ†’ Resolution โ†’ Review โ†’ Prevention

Key Concepts

Severity Levels

LevelDefinitionResponse
SEV1Critical โ€” Major customer impact, revenue lossAll hands, war room
SEV2High โ€” Significant degradation, workaround existsOn-call + escalation
SEV3Medium โ€” Minor impact, limited scopeOn-call handles
SEV4Low โ€” Minimal impact, cosmeticNext business day

Incident Roles

  • Incident Commander (IC) โ€” Coordinates response, makes decisions
  • Communications Lead โ€” Updates stakeholders, status page
  • Operations Lead โ€” Technical investigation and remediation
  • Scribe โ€” Documents timeline and actions

MTTX Metrics

  • MTTD โ€” Mean Time To Detect
  • MTTA โ€” Mean Time To Acknowledge
  • MTTR โ€” Mean Time To Resolve
  • MTBF โ€” Mean Time Between Failures

Topics

  • โ˜ On-call rotations
  • โ˜ Escalation paths
  • โ˜ War room protocols
  • โ˜ Status page management
  • โ˜ Customer communication
  • โ˜ Postmortem process
  • โ˜ Blameless culture
  • โ˜ Runbook development
  • โ˜ Incident tooling
  • โ˜ Chaos engineering

On-Call

Sustainable On-Call

  • No more than 25% of time in on-call
  • Maximum 2 incidents per shift
  • Compensatory time off after pages
  • Clear escalation when overwhelmed

On-Call Checklist

  • โ˜ Laptop and connectivity
  • โ˜ VPN access working
  • โ˜ Alert routing confirmed
  • โ˜ Runbooks accessible
  • โ˜ Escalation contacts known

Postmortems

Structure

  1. Summary โ€” What happened, impact, duration
  2. Timeline โ€” Detailed sequence of events
  3. Root Cause โ€” Why it happened (5 Whys)
  4. Impact โ€” Users affected, revenue lost
  5. Action Items โ€” Specific, assigned, time-bound
  6. Lessons Learned โ€” What we'll do differently

Blameless Culture

  • Focus on systems, not individuals
  • "What failed" not "who failed"
  • Assume good intentions
  • Share learnings widely

Runbooks

Good Runbook

  • Clear trigger (when to use)
  • Step-by-step actions
  • Expected outcomes at each step
  • Escalation criteria
  • Rollback procedures
  • Recently tested

Tools

ToolPurpose
PagerDutyAlerting and on-call
OpsgenieIncident management
StatuspageExternal communication
Jira/LinearAction item tracking
Slack/TeamsWar room coordination
Rootly/Incident.ioIncident automation

Anti-Patterns

  • Hero culture (one person fixes everything)
  • Blame-driven postmortems
  • Action items that never get done
  • Runbooks that don't work
  • Alert fatigue leading to ignored pages

Regulatory and standard mappings

Incident management controls

Mandatory incident reporting timelines

Integrated incident response workflow should accommodate multi-regime parallel reporting.

Reading

  • Google SRE Book: Chapters 12-15 (Effective Troubleshooting, Emergency Response, Postmortem Culture)
  • Incident Management for Operations (O'Reilly)

People-substrate cross-cluster

Incidents fire the nervous system. The operator's defense response under incident pressure is the lagging indicator of substrate. Blameless postmortems require real emotional maturity; without it, postmortems produce scapegoat-child dynamics at organizational scale.

  • Bridge essay: Interview Training as Applied Clinical Psychology -- Item 7 (the debrief room is a miniature dysfunctional family unless engineered against)
  • Developmental-position: Defense-Response Model -- which defense (fight / flight / freeze / fawn) the operator runs under incident pressure; the response is data, not character
  • Developmental-position: Freeze Cascade + Denial Cascade -- the failure cascades that produce MTTR blow-out
  • Psychology: Nervous-system regulation patterns -- communication shutdown under dysregulation = the MTTR-degrading mechanism at the human layer
  • Psychology: Real emotional maturity -- prerequisite for blameless postmortem to land; without it, "blameless" is performed while substrate runs blame
  • Anti-pattern (scapegoat): Scapegoat-child dynamics -- the team-level pattern that hijacks postmortems; one person absorbs the system's projection
  • Competency: Conflict Capacity + Emotional Regulation -- the dimensions incident commanders need; SEV1 reveals which dimension is under-developed

Atoms