"Security is a process, not a product."
What is Security in SRE?
Protecting systems, data, and users from unauthorized access, breaches, and attacks โ while maintaining reliability and velocity.
Core Principles
Defense in Depth
Multiple layers of security controls:
Network โ Infrastructure โ Application โ Data
Least Privilege
- Minimum permissions necessary
- Time-bound access
- Regular access reviews
Zero Trust
- Never trust, always verify
- Authenticate everything
- Encrypt everything
Shift Left
- Security early in development
- Automated security testing
- Developer education
Key Concepts
CIA Triad
- Confidentiality โ Only authorized access
- Integrity โ Data is accurate and unmodified
- Availability โ Systems accessible when needed
Attack Surface
- External endpoints
- Internal services
- Third-party integrations
- Supply chain (dependencies)
Threat Modeling
- What are we building?
- What can go wrong?
- What are we doing about it?
- Did we do a good job?
Topics
- โ Secrets management
- โ Identity and access management
- โ Network security
- โ Container security
- โ Supply chain security
- โ Vulnerability management
- โ Security monitoring
- โ Incident response
- โ Compliance automation
- โ Penetration testing
Secrets Management
Don't
- Secrets in code
- Secrets in environment variables (visible in process list)
- Secrets in CI/CD logs
- Shared secrets
Do
- Secrets in vault (HashiCorp Vault, AWS Secrets Manager)
- Dynamic/short-lived credentials
- Rotation automation
- Audit logging
Container Security
Image Security
- Minimal base images (distroless, alpine)
- No root user
- Scan for vulnerabilities
- Sign images
Runtime Security
- Read-only filesystem
- Drop capabilities
- Resource limits
- Network policies
Kubernetes Security
securityContext:
runAsNonRoot: true
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
Supply Chain Security
SLSA Framework (Levels 1-4)
- Source integrity
- Build integrity
- Provenance
- Common requirements
Dependency Management
- Pin versions
- Automated updates (Dependabot, Renovate)
- Vulnerability scanning
- SBOM generation
Tools
| Tool | Purpose |
|---|---|
| HashiCorp Vault | Secrets management |
| AWS Secrets Manager | Cloud secrets |
| Trivy | Container scanning |
| Snyk | Dependency scanning |
| Falco | Runtime security |
| OPA/Gatekeeper | Policy enforcement |
| SOPS | Encrypted secrets in Git |
| Teleport | Zero-trust access |
Security Monitoring
What to Monitor
- Authentication failures
- Authorization failures
- Privileged operations
- Data access patterns
- Network anomalies
- Configuration changes
Alerting
- Failed login attempts (brute force)
- Privilege escalation
- Unusual data access
- Policy violations
Anti-Patterns
- Security as afterthought
- Perimeter-only security
- Long-lived credentials
- Shared accounts
- No audit logging
- Security through obscurity
Compliance
Cross-sector InfoSec management
- ISO 27001 Cluster โ ISMS standard. Annex SL aligned. Certifiable. Cl 4-10 + 93 Annex A controls.
- SOC 2 Cluster โ AICPA TSC attestation. US procurement default.
- NIST CSF Cluster โ voluntary US framework. Six functions (Govern/Identify/Protect/Detect/Respond/Recover).
- BSI IT-Grundschutz Cluster โ DE national framework. ISO 27001 auf Basis IT-Grundschutz path.
- HITRUST Cluster โ healthcare-broadened. e1 / i1 / r2 certification levels.
EU regulatory
- GDPR Cluster โ Reg 2016/679. Personal data protection. โฌ20M / 4% turnover penalties.
- NIS2 Cluster โ Dir 2022/2555. Essential / important entities. โฌ10M / 2% turnover penalties.
- DORA Cluster โ Reg 2022/2554. Financial services ICT resilience. Applicable 17 Jan 2025.
- Cyber Resilience Act Cluster โ Reg 2024/2847. Product cybersecurity. Applicable late 2027.
- EU AI Act Cluster โ Reg 2024/1689. Risk-tier AI obligations. 2026-2027 ramp.
AI governance
- ISO 42001 Cluster โ AI Management System. Annex SL sibling to ISO 27001.
- NIST AI RMF Cluster โ voluntary US AI risk framework. GenAI Profile (NIST AI 600-1).
- OWASP LLM Top 10 Cluster โ operational LLM threat taxonomy (v2.0 / 2025).
- OECD AI Cluster โ international voluntary AI principles.
Threat intelligence
- MITRE Cluster โ ATT&CK + D3FEND + ATLAS frameworks.
Sector-specific
- PCI DSS Cluster โ Payment Card Industry Data Security Standard v4.0.
- TISAX Cluster โ German automotive supplier assurance.
- ISO 21434 Cluster โ Road vehicles cybersecurity engineering.
- UN R155 R156 Cluster โ Vehicle CSMS + SUMS type approval.
- FedRAMP CMMC Cluster โ US federal cybersecurity compliance.
- CSA CCM Cluster โ Cloud Security Alliance Cloud Controls Matrix.
Supply chain integrity
- SLSA SBOM Cluster โ Software supply chain (SLSA levels + SBOM via SPDX / CycloneDX). Maps to ISO 27001 A.5.21 + A.8.30.
Continuity and resilience
- ISO 22301 Cluster โ Business Continuity Management Systems.
Governance and process
- COBIT Cluster โ IT governance framework.
- ITIL Cluster โ IT service management framework.
- TOGAF Cluster โ enterprise architecture framework.
Automation
- Policy as code
- Continuous compliance
- Evidence collection
- Drift detection
Reading
- Google SRE Book: Chapter 9 (Simplicity)
- Building Secure & Reliable Systems (Google)
- OWASP guidelines
Atoms
Published expressions
- "Shifting left security" is a misnomer that needs to die (9-part series) -- the long-form argument behind this pillar's "security is a process, not a product" stance. Start at Part 1 of 9; the series chains prev/next through Part 9 of 9.
Homelab worked example
Stefan's home LAN is a live SRE testbed. Security-relevant pieces (2026-05-15 hardening pass):
- SSH hardening on
[host]and[host]:PasswordAuthentication no,PubkeyAuthentication yes, passwordless sudo via/etc/sudoers.d/slaine-nopasswd. Drop-in01-hardening.confloads before cloud-init defaults. See SRE/homelab addendum ยง3 andreference_pihole_hardening(memory). - Single-key SPoF: one SSH key on the Mac unlocks
[host]+[host]+[host]with passwordless sudo, chains to AWS via[host]. Memory:reference_slaine_lan_master_key. Mitigation candidates: hardware-backed key (YubiKey resident), passphrase + agent-forwarding discipline. - Plaintext-secret incident + remediation (2026-05-15):
~/Ni-adeleted; vodafone DSL credentials still at~/.config/vodafone-dsl/credentialsmode 600 (not key-rotatable). Migration target: SOPS or age-encrypted vault. - Zero-trust position (addendum ยง4 hardening pass): tailscale node identity as authorisation primitive, not LAN location. Pi-hole web UI is currently anonymous on both nodes โ closes when admin password set + tailscale ACLs scoped per device.
- Break-glass design (parked): pre-baked AMI in second AWS region, MFA-locked launch, CloudTrail-tagged recovery events. Quarterly boot-and-verify drill required (else untested = unreliable).
- Headscale single point of failure: if everything routes through Headscale ACLs, Headscale outage = ops outage. HA pair in same region only partially mitigates.
- WD MyCloud
[host]: end-of-life 2015 firmware, exposed SSH (ssh-rsa + ssh-dss legacy ciphers), NFS, SMB, AFP. Highest-risk surface on the LAN. Audit-and-decom listed in SoT ยง7 open items.
People-substrate cross-cluster
Security at the people layer = psychological safety + boundary integrity + structural defenses against the team-level failure modes. Fawning is the human auth-bypass (yes-by-default = compromised input validation). Scapegoat dynamics are the team-level incident-response failure that targets a projection surface. Narcissism is the persistent-threat actor at organisational scale.
- Bridge essay: Interview Training as Applied Clinical Psychology -- Item 7 (debrief room as miniature dysfunctional family)
- Competency: Psychological Safety dimension -- the team-level confidentiality + integrity property; people speak up = honest input
- Competency: Boundary Integrity dimension -- least-privilege at the relational layer
- Psychology: Fawning -- the auth-bypass failure mode; nervous-system-default yes is the human equivalent of permissive ACLs
- Psychology: Narcissism and relational abuse patterns -- the APT at human scale; entropy-management at the team's expense
- Psychology: Scapegoat-child dynamics -- the team-level incident-response failure mode; analogous to mis-routing a security alert to the wrong service
- Developmental-position: Social-Defense Band -- diagnostic for whether reports are running fawn / fight / flee around the leader; the security posture of the team