"Reliability is the most important feature."
What is Reliability?
The ability of a system to perform its intended function under stated conditions for a specified period of time.
Key Concepts
Service Level Indicators (SLIs)
Quantitative measures of service behavior:
- Availability โ % of successful requests
- Latency โ Response time distribution (p50, p95, p99)
- Throughput โ Requests per second
- Error rate โ % of failed requests
- Freshness โ Data staleness
Service Level Objectives (SLOs)
Target values for SLIs:
Availability SLO: 99.9% of requests successful over 30 days
Latency SLO: p99 < 200ms over 30 days
Service Level Agreements (SLAs)
Contractual commitments with consequences:
- SLA = SLO + Business consequences (refunds, penalties)
- SLAs should be less aggressive than internal SLOs
Error Budgets
The inverse of reliability โ how much failure is allowed:
99.9% SLO = 0.1% error budget = 43.2 minutes/month downtime allowed
99.99% SLO = 0.01% error budget = 4.32 minutes/month downtime allowed
Topics
- โ Defining meaningful SLIs
- โ Setting realistic SLOs
- โ Error budget policies
- โ Reliability vs. velocity tradeoffs
- โ Cascading failures
- โ Redundancy and replication
- โ Graceful degradation
- โ Circuit breakers
- โ Retry strategies with backoff
- โ Timeouts and deadlines
Patterns
- N+1 redundancy
- Active-passive failover
- Active-active clustering
- Geographic distribution
- Bulkhead isolation
Anti-Patterns
- Assuming the network is reliable
- Single points of failure
- Cascading timeouts
- Retry storms
- Unbounded queues
Tools
| Tool | Purpose |
|---|---|
| Prometheus | SLI measurement |
| Grafana | SLO dashboards |
| Sloth | SLO generator for Prometheus |
| OpenSLO | Vendor-neutral SLO specification |
Reading
- Google SRE Book: Chapters 3-4 (Embracing Risk, Service Level Objectives)
- The Site Reliability Workbook: Chapter 2 (Implementing SLOs)
Regulatory and control mappings
- ITIL 4 Practices Service Level Management practice.
- ISO 27001 Annex A.5 Organizational Controls A.5.29-A.5.30 disruption / ICT readiness for BC.
- ISO 22301 Clause Structure and Key Concepts BIA + RTO + RPO + MBCO discipline.
- DORA ICT Risk Management Art 11 response and recovery.
- NIST CSF Core Functions RECOVER function.
Atoms
Homelab worked example
Stefan's home LAN is a live SRE testbed. Reliability-relevant pieces:
- Pi-hole HA pair (
[host]MASTER +[host]BACKUP + VIP[host]via keepalived VRRP) โ addresses the historical Pi-hole DNS SPoF that took the whole LAN down on any single-host outage. Failover ~4-5s on tested scenarios. See SRE/homelab addendum ยง1 andconstraint_pihole_dns_chain(memory). Outstanding: gravity.db replication (gravity-sync v4). - Headscale HA on AWS EC2 (
headscale-primary+headscale-standbywith EIPs, terraform-managed). See work-kb SoT ยง5. - Migration sequencing rule (homelab, agreed 2026-05-15): wait โ cutover โ steady state โ THEN harden. Reliability work before security work; don't pile changes mid-flight.
- Single-NAT vs double-NAT (DrayTek bridge migration ยง6): conntrack exhaustion on the consumer-grade WAN device is a Pi-hole-orthogonal availability failure; documented symptom + planned fix.
People-substrate cross-cluster
Reliability has a people-substrate dual. The leader's regulation under load is an SLI; trauma-substrate adaptations break the same way under-engineered systems break (cascading failure, hero patterns, single points of failure).
- Bridge essay: Interview Training as Applied Clinical Psychology
- Psychology: Real emotional maturity -- the structural reliability property of the operator; performance of maturity is the apparent layer that fails under load
- Psychology: Lack of accountability predicts relationship failure -- Gottman empirical anchor; inability to own a mistake is the relational-SLO breach predictor
- Psychology: Nervous-system regulation patterns -- dysregulated operator = unreliable response under incident; the regulation layer IS the reliability layer
- Anti-pattern (hero culture): Caretaker syndrome + Hyper-independence -- the human SPoF; "I just do it myself" is the same failure pattern as a non-redundant database
- Anti-pattern (SPoF): Ownership psychology over-indexed = the no-redundancy condition the SRE Reliability discipline exists to prevent