3. Observability

Vault note, not reviewed against the source. Written in the knowledge vault on 2026-04-25 by models working with Stefan Coetzee and published as it stands, with private addresses, e-mail addresses and an employer name redacted. Check claims against the primary source before relying on them.

"You can't fix what you can't see."

What is Observability?

The ability to understand a system's internal state by examining its external outputs โ€” without deploying new code.

Three Pillars

1. Metrics

Numeric measurements over time:

  • Counters โ€” Cumulative values (requests_total)
  • Gauges โ€” Current values (temperature, queue_size)
  • Histograms โ€” Distribution of values (latency_bucket)
  • Summaries โ€” Quantiles (p50, p95, p99)

2. Logs

Discrete events with context:

  • Structured (JSON) > Unstructured
  • Include: timestamp, level, service, trace_id, message, context
  • Sampling for high-volume services

3. Traces

Request flow across services:

  • Trace = collection of spans
  • Span = single operation with timing
  • Propagate context (trace_id, span_id) across services

Key Concepts

RED Method (Request-driven)

  • Rate โ€” Requests per second
  • Errors โ€” Failed requests per second
  • Duration โ€” Latency distribution

USE Method (Resource-driven)

  • Utilization โ€” % time resource is busy
  • Saturation โ€” Work queued waiting
  • Errors โ€” Error count

Four Golden Signals

  1. Latency
  2. Traffic
  3. Errors
  4. Saturation

Topics

  • โ˜ Metric naming conventions
  • โ˜ Cardinality management
  • โ˜ Log aggregation pipelines
  • โ˜ Distributed tracing implementation
  • โ˜ Correlation (metrics โ†” logs โ†” traces)
  • โ˜ Alerting strategies
  • โ˜ Dashboard design
  • โ˜ Runbook integration
  • โ˜ Cost of observability
  • โ˜ Sampling strategies

Alerting

Good Alerts

  • Actionable โ€” Someone needs to do something
  • Urgent โ€” It can't wait
  • Symptom-based โ€” Users are affected
  • Documented โ€” Runbook linked

Bad Alerts

  • Noisy โ€” Fires too often, gets ignored
  • Cause-based โ€” CPU high but no user impact
  • Ambiguous โ€” What should I do?

Alert Hierarchy

Page (wake someone up)
  โ†’ Ticket (fix during business hours)
    โ†’ Log (investigate when time permits)

Tools

CategoryTools
MetricsPrometheus, Datadog, CloudWatch
LogsELK, Loki, Splunk
TracesJaeger, Tempo, Zipkin, X-Ray
VisualizationGrafana, Kibana
AlertingAlertmanager, PagerDuty, Opsgenie

Anti-Patterns

  • Alert fatigue (too many alerts)
  • Vanity metrics (dashboards no one uses)
  • High cardinality labels
  • Missing context in logs
  • No correlation between signals

Reading

  • Google SRE Book: Chapters 6, 10 (Monitoring, Alerting)
  • Distributed Systems Observability (O'Reilly)

Regulatory and control mappings

Homelab worked example

Stefan's home LAN is a live SRE testbed. Observability-relevant pieces:

  • Prometheus + Grafana + Alertmanager stack on Synology [host] (/volume1/docker/monitoring/). Grafana on :3030, Prometheus on :9090. See work-kb SoT ยง3 service inventory.
  • Exporters: node-exporter on [host] + container on [host], pihole-exporter :9617, blackbox-exporter :9115, speedtest-exporter :9696, unifi-poller :9130 (pulls UCG-Max controller metrics).
  • NetBox at http://[private IP]:8000 โ€” structured device + IP inventory. Drives the SoT atom's host-inventory table via ~/code/scripts/regen-network-atom.py. Combined with UniFi controller pull (regen-unifi-atom.py) for live state.
  • Central rsyslog collector at [private IP]:514 UDP/TCP. Logs land in /var/log/remote/<ip>/. RADIUS deferred. Memory: reference_rsyslog_161.
  • Cardinality discipline: NetBox device records are the cardinality cap for per-host metrics; ARP-discovered ghost clients flagged as apple-iot-placeholder-1 rather than minted as net-new devices each scrape.
  • Audit-log gap (per homelab addendum hardening backlog ยง4): sshd + sudo + headscale + pi-hole logs are not yet shipped off-host. Local-only logs equal no audit when LAN compromised.

People-substrate cross-cluster

Observability has a people-substrate dual: reading what's actually happening in the team (and in yourself) is the same skill as reading what's happening in a system. The three SRE pillars (metrics, logs, traces) map to substrate / apparent / assigned in the developmental-position frame. The trauma filter is the observability bias the operator has to debias against.

  • Bridge essay: Interview Training as Applied Clinical Psychology -- Item 5 (cognitive distortions = observability bias) + Item 11 (post-interview capture = session-note discipline)
  • Developmental-position: Three-Layer Position Model -- substrate / apparent / assigned is the observability stack for humans; most "observation" stays on apparent
  • Developmental-position: The Trauma Filter -- the observability bias every operator carries; rubrics + structural defenses are the procedural counter
  • Competency: Self-Awareness dimension -- the operator's own observability stack; without it, all incident signals get filtered through ego-protection
  • Psychology: Cognition and influence -- DMN identity decoupling; default-mode network produces noise that masquerades as signal
  • Psychology: Amygdala as smoke detector -- the legacy alerting layer that fires on pattern-match, not on actual threat; analogous to alert-fatigue mechanism in monitoring

Atoms