1. Reliability

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-04-25 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

"Reliability is the most important feature."

What is Reliability?

The ability of a system to perform its intended function under stated conditions for a specified period of time.

Key Concepts

Service Level Indicators (SLIs)

Quantitative measures of service behavior:

  • Availability โ€” % of successful requests
  • Latency โ€” Response time distribution (p50, p95, p99)
  • Throughput โ€” Requests per second
  • Error rate โ€” % of failed requests
  • Freshness โ€” Data staleness

Service Level Objectives (SLOs)

Target values for SLIs:

Availability SLO: 99.9% of requests successful over 30 days
Latency SLO: p99 < 200ms over 30 days

Service Level Agreements (SLAs)

Contractual commitments with consequences:

  • SLA = SLO + Business consequences (refunds, penalties)
  • SLAs should be less aggressive than internal SLOs

Error Budgets

The inverse of reliability โ€” how much failure is allowed:

99.9% SLO = 0.1% error budget = 43.2 minutes/month downtime allowed
99.99% SLO = 0.01% error budget = 4.32 minutes/month downtime allowed

Topics

  • โ˜ Defining meaningful SLIs
  • โ˜ Setting realistic SLOs
  • โ˜ Error budget policies
  • โ˜ Reliability vs. velocity tradeoffs
  • โ˜ Cascading failures
  • โ˜ Redundancy and replication
  • โ˜ Graceful degradation
  • โ˜ Circuit breakers
  • โ˜ Retry strategies with backoff
  • โ˜ Timeouts and deadlines

Patterns

  • N+1 redundancy
  • Active-passive failover
  • Active-active clustering
  • Geographic distribution
  • Bulkhead isolation

Anti-Patterns

  • Assuming the network is reliable
  • Single points of failure
  • Cascading timeouts
  • Retry storms
  • Unbounded queues

Tools

ToolPurpose
PrometheusSLI measurement
GrafanaSLO dashboards
SlothSLO generator for Prometheus
OpenSLOVendor-neutral SLO specification

Reading

  • Google SRE Book: Chapters 3-4 (Embracing Risk, Service Level Objectives)
  • The Site Reliability Workbook: Chapter 2 (Implementing SLOs)

Regulatory and control mappings

Atoms

Homelab worked example

Stefan's home LAN is a live SRE testbed. Reliability-relevant pieces:

  • Pi-hole HA pair ([host] MASTER + [host] BACKUP + VIP [host] via keepalived VRRP) โ€” addresses the historical Pi-hole DNS SPoF that took the whole LAN down on any single-host outage. Failover ~4-5s on tested scenarios. See SRE/homelab addendum ยง1 and constraint_pihole_dns_chain (memory). Outstanding: gravity.db replication (gravity-sync v4).
  • Headscale HA on AWS EC2 (headscale-primary + headscale-standby with EIPs, terraform-managed). See work-kb SoT ยง5.
  • Migration sequencing rule (homelab, agreed 2026-05-15): wait โ†’ cutover โ†’ steady state โ†’ THEN harden. Reliability work before security work; don't pile changes mid-flight.
  • Single-NAT vs double-NAT (DrayTek bridge migration ยง6): conntrack exhaustion on the consumer-grade WAN device is a Pi-hole-orthogonal availability failure; documented symptom + planned fix.

People-substrate cross-cluster

Reliability has a people-substrate dual. The leader's regulation under load is an SLI; trauma-substrate adaptations break the same way under-engineered systems break (cascading failure, hero patterns, single points of failure).

  • Bridge essay: Interview Training as Applied Clinical Psychology
  • Psychology: Real emotional maturity -- the structural reliability property of the operator; performance of maturity is the apparent layer that fails under load
  • Psychology: Lack of accountability predicts relationship failure -- Gottman empirical anchor; inability to own a mistake is the relational-SLO breach predictor
  • Psychology: Nervous-system regulation patterns -- dysregulated operator = unreliable response under incident; the regulation layer IS the reliability layer
  • Anti-pattern (hero culture): Caretaker syndrome + Hyper-independence -- the human SPoF; "I just do it myself" is the same failure pattern as a non-redundant database
  • Anti-pattern (SPoF): Ownership psychology over-indexed = the no-redundancy condition the SRE Reliability discipline exists to prevent