The principle
Describe the desired state of infrastructure and configuration in files under version control, review changes like code, and apply them with a tool that reaches the same end state however often it runs. The repository is then the system's claim about itself, with a history of who changed what and why. A change made by hand on a server is drift, and drift tends to surface in the next outage.
Source: Infrastructure as Code and Idempotence as the IaC Invariant; the configuration management section of Google's Release Engineering.
On this platform
In git, public:
- The sites. Every page, generator and stylesheet of machinebehavior.io (repository).
- The portal's data. services/*.yml, incidents/*.yml, the navigation in
site/nav.yml, and the decision records. - The gate. Its code, the requirements under test and the rule tiers (
conformity/requirements.json,conformity/site-tier.json), and its run history inconformity/latest.json(Conformity gate). - The observability stack. The compose file, the Prometheus rules with their tests, the Loki ruler rule, the collector configuration and the dashboards generated by
grafana/build.py(agent-observability, Dashboards as code). Grafana loads the dashboards read-only, so the UI holds no edits of its own.
Not in git:
- Which dashboards are public. The share record is created by an API call on the home server, and its token exists only in Grafana (Share a dashboard).
- The hosts. The home server and the cloud proxy under the stack are not described in a public repository; this review did not check whether their setup could be rebuilt from code.
- GitHub settings. Labels, milestones and the Pages configuration live in GitHub. The Ticket board page documents the labels; no file in the repository defines them.
- The AWS budget. It is set in the account, and its amount is not published (Budgets and alerts).
State
Partial. Everything a reader meets, and the rules that watch it, is in git with its history. The layers below (hosts, sharing, repository settings) are configured by hand, and a rebuild of them would depend on runbooks and memory.
Services that name it
In a company of 10 to 200 people
- Put the cloud accounts, DNS, CI settings and monitoring rules in code first; they are the parts nobody can rebuild from memory during an outage.
- Turn off write access in the cloud console for daily work, and treat a console change as an incident to backfill into code.
- Run the plan or diff step on a schedule and alert on drift.
- Keep the service catalog, the runbooks and the decision records in the same repositories as the code they describe.