Operations and SRE
Running software and agents in production. Incidents written blameless, runbooks, error budgets, toil, and the operations practice that carries over to language models.
1 research page, 6 Inside pages, 119 docs pages, 12 posts, 5 tickets
Research 1
- Chaos engineering for behaviour
Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it.
Inside 6
- Mission Control
The live operations view of the platform behind machinebehavior.io on one screen: gate and deploy state of the three sites, the deploy feed, open incidents, the conformity gate result and the public Grafana dashboards.
- Status
Status of the 12 services behind machinebehavior.io, read from the gate records and the deploy feed, and the incident history with the stages Investigating, Identified, Monitoring and Resolved.
- Services
The service catalog of the platform behind machinebehavior.io: 12 services with owner, tier, lifecycle, SLO, dashboard, runbooks, docs, repository and dependencies, built from YAML in the repository.
- Board
The open work on machinebehavior.io and the platform behind it as a board: backlog, ready, in progress, blocked and done. GitHub Issues on the public repository is the system of record; the board is a snapshot taken at each deploy.
- Deploy job
The second job of the Conformity workflow. It runs only after the checks pass, refreshes the map and the search index, writes the deploy feed and the board snapshot, and deploys the site to GitHub Pages.
- machinebehavior.io
The research site and the home of Inside: claims, experiments, objections, case files, the map, the docs, the board and this catalog. Hand-written HTML and Python generators, served by GitHub Pages.
Docs 119
Engineering 4
- Gate blocked a deployThe conformity job failed, so the deploy did not run and the previous build is still live; find the failing check, fix the page, push again.
- Link check flags a template stringThe gate reads every href in the page source, including ones inside JavaScript template strings; build such links with a.href in code.
- Push rejected after a deployThe conformity bot commits after every run, so main moves without you; rebase on it and push again.
- RunbooksStep-by-step fixes for the failures seen so far in the build and deploy of the sites.
Observability 6
- Alerts and SLOsThe recording rules, service level objectives and 17 alerts that Prometheus and the Loki ruler evaluate, with thresholds from the rule files.
- Backfilled runs missing in LokiWhat to do when the Website deploys dashboard shows no older runs after the deploy-exporter backfills from GitHub Actions.
- Counter resets from parallel sessionsWhy Claude Code's cost counters in Prometheus report spend in the millions of USD, and how the stack reads spend and tokens from per-request events in Loki.
- RunbooksStep-by-step fixes for the known traps in the observability stack: missing backfill in Loki, empty stat panels, Claude Code counter resets and sharing a dashboard.
- Share a dashboard publiclyHow to build a variable-free public cut of a dashboard, provision it, and turn on Grafana public sharing so it loads at grafana.scoetzee.de.
- Stat panel shows No dataWhy Grafana stat panels on young series show No data over long time ranges, and how build.py switches them to instant queries.
SRE Handbook 109
- SRE HandbookSite reliability engineering from the knowledge vault: the manifesto, ten pillars, patterns, runbooks, tools and incident records.
- 1. Reliability"Reliability is the most important.
- 10. Toil Reduction"If a human is doing something a computer could do, that's a.
- 2. Scalability"Scale is not a feature you can bolt on.
- 3. Observability"You can't fix what you can't.
- 4. Incident Management"It's not about preventing all failures.
- 5. Infrastructure as Code"If it's not in Git, it doesn't.
- 6. CI/CD & Deployment"If it hurts, do it more.
- 7. Performance"Premature optimization is the root of all evil.
- 8. Security"Security is a process, not a.
All 109 pages in SRE Handbook
- 9. Cost Optimization"There is no cloud. It's just someone else's computer, and they're billing you for.
- AI Agents are Ops WorkManaging the lifecycle of an AI agent in production is operational engineering, not developer engineering.
- AI agents are ops workAn AI agent in production is a service and needs what services need: owners, SLOs, paging, runbooks, cost control and canaries. Here the agent sessions are a catalogued service that pushes through the same gate, with metered spend and an incident record; they have no SLOs and no canary set.
- Apple Silicon vs Desktop GPU for InferenceA16 Neural Engine at 17 TOPS with unified memory beats a GTX 1650 4GB at local LLM inference.
- Authenticated Cookie Jar for Gated PostsAn exported cookies.txt lets yt-dlp list and fetch gated posts.
- Back Up SQLite With the Backup CommandUse sqlite3 .backup, or checkpoint the WAL first, to get a restore point that holds every write.
- Blameless Postmortem DisciplineA blameless postmortem assumes the engineer behaved rationally given the information they had.
- blind-search-ratchet-minecraft-beats-itselfA simulation of Minecraft with no player let random mob behaviour kill the Ender Dragon after about 1.975 billion simulated years.
- BM25 Hybrid Retrieval for Graph-RAGBM25 is a 1990s ranking function that's still the backbone of serious retrieval.
- Bracket Pattern Process CheckUse pgrep -f "[b]atchtranscribe.py" or the ps | awk form.
- Certificate RotationReplace TLS certificates before expiry.
- Claude Data Export for Graph IngestionPipeline for pulling Claude conversation history into a graph-RAG vault.
- Compaction is the New OOMLLM context exhaustion is the operational hazard of AI-era systems.
- Compare Videos Found Against the In-App CountThe owner's in-app count minus videosfound measures the gap.
- Compare Videos Found Run to RunFirst fix for the short scrape: treat any drop in videosfound between runs as suspect.
- Context Window Sizes and Effective RangeAdvertised context windows are marketing.
- Cost per unit of workCost is an operational signal read per unit of output, metered like latency and alerted like errors. Here a FinOps space meters each component, prices spend per call, commit and deploy, and sets alert thresholds from data; only the AWS budget e-mail reaches a person.
- CPU Cache Hierarchy and Speculative ExecutionMain memory is the source of truth.
- Credential RotationReplace secrets used by services and humans.
- Database FailoverPromote a replica to primary when the current primary is unhealthy.
- Defense in Depth and Trust BoundariesNo single security control holds against a determined adversary.
- Differential Test in the Same MinutesRun the failing path and its nearest working variants side by side, within the same few minutes, changing one factor at a time.
- Dolt Versioned Database for Task StateMySQL-compatible database with Git-like versioning.
- Eliminate Concurrent Saves Before Blaming the ScraperAny 'scrape found N, collection holds M' gap must first rule out that the owner saved something mid-run.
- Eliminate toilToil is manual, repetitive, automatable work that grows with load and leaves nothing behind; SRE caps it so engineering time remains. Here generators and the deploy job removed most of it; a rebase before every push, hand deploys of the observability stack and the screenshot retakes remain, and nobody measures the share.
- Enumeration Retry With Unauthenticated FallbackEnumeration tries the cookie jar twice, then unauthenticated mode three times with backoff, and logs which path succeeded.
- Error Budgets as Reliability CurrencyAn error budget is the only mechanism that makes the reliability-vs-velocity trade-off concrete.
- Every change passes the same gateRelease engineering makes every change go through one repeatable, recorded path, with builds that give the same output wherever they run. Here one workflow gates and deploys three sites and records every run; the generated pages are built on the author's machine, and the observability stack has no pipeline.
- Every incident ends in a recordEach incident is written up without blame, with impact, timeline, cause and follow-up work, and the lesson becomes a runbook or a fix. Here four incident records sit on the status page with stages, write-ups and tickets; detection times, a template and an index are missing.
- Expired Bot-Manager Cookie Returns Empty Bodyyt-dlp sends expired entries from a Netscape cookie jar as they are.
- File Copy of a WAL-Mode Database Misses Recent WritesA SQLite database in WAL mode keeps recent writes in the -wal file.
- Filter Expired Cookies Before Each Runcookieargs() writes a filtered copy of the cookie jar with expired entries dropped, logs each drop, and passes the filtered file to yt-dlp.
- Generate-and-Test Over Reason-From-ModelWhen a system has enough internal constraint, reasoning about which change should work reliably loses to generating many changes and testing which ones do.
- Graph-RAG over Flat RAG for Operational KnowledgeOperational knowledge is structurally relational: an incident references a service references a deployment references a runbook references a metric references an SLO.
- Hand-Check Output Against the Primary SourceCompare a sample of generated summaries or extracted figures with the source they came from.
- Handle-Based Ad Filter Matches Artist AccountsThe creator rule \.(de|official)$ was written for brand and shop accounts.
- Headscale Mesh VPN for Data SovereigntySelf-hosted open-source Tailscale control plane with embedded DERP relay.
- Horizontal vs Vertical ScalingThe choice is not "which is better", it is "where does the bottleneck actually live, and which dimension can absorb it." Vertical scaling is bounded by hardware.
- Idempotence as the IaC InvariantThe property that makes Infrastructure as Code work is not "code that builds infrastructure." It is idempotence: applying the same configuration to the same target produces the same end-state, no matter how many times you run it or what the prior state was.
- If it is not in git, it does not existInfrastructure and configuration are declared in version control and applied the same way every time, so the repository describes the running system. Here the sites, the catalog, the incidents, the gate rules, the alert rules and the dashboards are in git; the hosts, the dashboard sharing and the GitHub settings are not.
- Incident 2026-02-27 Dead Collection Short LinksTwo collections stopped scraping because their stored short links died.
- Incident 2026-07 Summaries Inverted IronyOne sarcastic recommendation was summarised as a sincere warning, one summary named the wrong person, one was selectively pessimistic against its source.
- Incident 2026-07-14 Backup Missed WAL WritesA pre-change backup copied the main database file only, while the database was in WAL mode with a multi-megabyte WAL file.
- Incident 2026-07-14 Short Scrape Looked Like a Quiet FeedSix runs in a row reported 19 videos found; a later run reported 32, and the owner's own count was closer than the database's.
- Incident 2026-07-19 Pending Status on Finished VideosTwelve videos sat at statuspending while holding full transcripts and summaries.
- Incident 2026-07-21 Ad Filter Removed Artist VideosTen videos from artists with .official handles were rejected by the ad filter before transcription and looked like unprocessed backlog.
- Incident 2026-08-08 Empty-Transcript Videos Invisible to ReviewA reel made of on-screen text returned no transcript.
- Incident 2026-08-08 False Enumeration Bug FindingA scrape found 46 videos and a re-enumeration minutes later found 48.
- Incident 2026-09-11 Three-Morning Cron CrashThe 06:15 cron died three mornings running with CalledProcessError.
- Incident 2026-09-13 Ghost Still-Running JobA process check over SSH reported the batch job as still running after it had exited.
- Incident 2026-09-16 Server Cannot List the CollectionFrom 2026-09-16 the server could not enumerate the collection from any client or egress, while the laptop could with the same yt-dlp version through the same egress.
- IncidentsIncident, cause and mitigation records from the vault, written up from trap files and session logs. Each record names its source.
- legacy-restoration-as-sre-craftA man walks you through the one and only airworthy Lockheed C-121 Constellation, 10,000 horsepower, "not a straight line on her," MacArthur's bar in the aft lounge, a veteran of the Berlin Airlift and NASA.
- LLM as Software-Defined CPUAn LLM is a CPU that shipped without a memory management unit.
- ManifestoA practical definition of Site Reliability Engineering, outside any strict definitions made by Google or any other company.
- Memory Architecture L0-L4A five-layer memory stack for AI coding agents: L0 context, L1 beads, L2 memories, L3 vault, L4 Hivemind.
- Monitor symptoms, page on user painMonitoring should say what is broken for users before it says why, and a page should reach a person only when they must act. Here burn-rate and degraded-search alerts read what users get, five dashboards are public, and no alert is delivered.
- No-Transcript Bucket in the Review ToolPlanned: have the review tool surface rows with a NULL review and an empty transcript as their own 'needs a manual look' bucket, so they stop disappearing whether or not OCR ever lands.
- OpenCode Self-Hosted LLM ConfigurationOpenCode connects to any self-hosted LLM that exposes an OpenAI-compatible API via the @ai-sdk/openai-compatible provider.
- Operations work done as softwareSRE is operational knowledge plus software engineering, with automation in place of manual steps and tools in place of one-off scripts. Here generators build the catalog, status page, docs, map, board and changelog from sources in git, and each build checks its input.
- PatternsReusable architecture patterns for reliable, scalable systems.
- Personal Digital Twin ArchitecturePer-person vault (private, IP-owned by the individual) plus a curated public projection (expertise, decisions, communication patterns).
- Principles in practiceFourteen SRE principles, from Stefan Coetzee's framework with Google's SRE book as the reference, each tied to where this platform applies it, with its state (in place, partial or gap) and the open work.
- Probe Pending Rows Before Calling Them BacklogA pending count is unclassified until the rows are looked at.
- Probe the Artifact Before Reasoning From CountsWhen a count looks wrong, fetch one real item and read the error.
- Probing Microarchitecture for Vulnerability DiscoveryCVE discovery on CPUs follows a repeatable methodology: probe the gap between what the architecture claims and what the implementation actually does.
- Process Check Matches Its Own Shellpgrep -f <pattern> run over SSH matches the remote shell whose command line contains the pattern.
- Progressive Delivery PatternsBig-bang deploys treat every release as a coin flip on the entire user base.
- Publish Date Treated As Time In CollectionA video's publish date says nothing about when it was saved to a collection.
- Pull the Full Row List After Any ScrapeDo not take the review tool's count as the batch size.
- Re-Scrape Relinks Processed Rows As PendingRe-scraping a collection re-links videos that were already processed and sets them back to statuspending without flipping them to complete again.
- Review Tool Lists Only Rows With a Transcriptreviewstatus.py reports on rows where length(transcript) > 0.
- Risk is a budgetA target of 100% reliability costs more than users can notice; the gap between the target and 100% is an error budget for change. Here tiers set the order of response and the Local LLM has burn-rate alerts, but no policy says what happens when a budget is spent.
- Rollback DeploymentRevert production to the last known good state.
- RunbooksOperational procedures for common scenarios.
- Sensitivity-Gated Posts Vanish From Unauthenticated ListingTikTok refuses to serve age-restricted or sensitivity-flagged posts to an unauthenticated client.
- Service level objectivesAn SLO is a target for a measured indicator of what users get, and the basis for alerts and error budgets. Here one of twelve services has SLOs, the Local LLM; the five tier-1 services, which readers meet or which decide whether a deploy goes out, have none.
- Service Outage ResponseGeneric triage runbook for production service unavailability.
- Short Link Redirects to the HomepageStored vm.tiktok.com short links for two collections now redirect to the site homepage.
- SimplicityEvery line of code and every moving part is a liability, so prefer boring technology, small interfaces and deleted code. Here three static sites, standard-library Python, no third-party requests on load and one navigation source keep the platform small enough for one operator.
- Single Video URL As the Collection ArgumentA single video URL works as the --collection-url argument, so individual videos can be reprocessed without a working collection link.
- Small Local Model Misreads Ironic RegisterA 7B local summarizer renders sarcasm as sincere statement, swaps subjects, or turns selectively pessimistic against its own source.
- Someone carries the pagerThe test for SRE work is standby, a person who is told when production breaks and acts. Here 17 alert rules reach nobody, the gate is the one control that acts at any hour, and three of four incidents were found by someone working at the time.
- SRE as Truth Verified WorkingThe acronym is incidental.
- SRE is a role inside OpsSRE is Ops with software engineering applied, the outer ring of a concentric model, with FinOps, SecOps and DevOps as overlaps. Here one operator holds every ring, agent sessions do the work, and the catalog names an owner for each service.
- Summaries Are a Triage IndexUse summaries to decide what to read, never as evidence.
- Swallowed Stderr Hides the Causesubprocess.run(..., checkTrue, captureoutputTrue) captures the child's stderr and raises CalledProcessError.
- Symptoms over Causes for AlertingPage when users are affected.
- tencent-agent-memory-four-tierAn open-source, fully-local memory layer that gives an agent human-like long-term recall, so it stops starting from scratch every session.
- The Disappearing Full-Stack Ops EngineerThe concentric Ops model produces engineers at every ring (SysAdmin, Cloud, Platform, SRE).
- The ten pillarsThe ten pillars of the SRE framework, one page each: reliability, scalability, observability, incident management, infrastructure as code, CI/CD and deployment, performance, security, cost optimisation and toil reduction.
- Toil vs Engineering and the 50 Percent RuleSRE without a toil cap drifts into operations.
- ToolsTool-specific guides and configurations.
- Transcriber Incident RegisterEvery failure of the video transcriber pipeline since February 2026, as typed atoms.
- Truth, verified, workingAn operational claim must be true, backed by a probe in its time window, and about a system that works. The gate, the status page and the review labels apply it here; numbers written into the docs still go stale without a probe.
- Unit Economics of InfrastructureTotal cloud spend is the wrong metric.
- USE Method for Resource SaturationWhen a system is slow and you don't know why, ask three questions of every resource it depends on: how busy is it, what is queued waiting for it, and is it returning errors.
- vault-search-fastembed-migrationHow to move a vault-search install off Ollama-based embeddings onto in-process fastembed, and how to replicate the change on another machine.
- Verify Transcript Length Never StatusTreat status as a report. Check length(transcript) and length(summary) per row.
Posts 12
- Knowledge Infrastructure for LLMsA language model can only work from what someone wrote down and put in front of it. Most companies that bought AI maintain their code and little else.
- Chaos Engineering for Behaviour
- The Stance Layer Is Still ToilA hook can block a word on every reply. Nothing I run can yet catch a model before it folds, only after.
- Which LLM User Are We Talking About?Advice about language models depends on which model, on whose hardware, inside how much harness
- They Trained Out the Board Edit. The Cheating Moved.A reading of the Goodhart Labs chess honeypot, three additions to the design, and a dated prediction.
- An RCA on ClaudishWhere Claude's writing style came from: text written to be heard, read in silence
- Compaction Is the New OOMMost agent harnesses ship without swap.
- The Track: The Drivers Never Buy ItEveryone got a race car, but no track yet.
- Success Is the Engine RunningThe modern world was built on explosions: contained, timed and measured ones. That is what an engine is. Language model output is the fire; the harness, the hook and the far-end gauge are the engine.
- What Operations Already Knows About Running AgentsError budgets, reconciliation loops, separation of duties and recovery over prevention: four operations practices that fit agent work almost line for line.
- The Trap File Is Longer Than the Instruction FileSeven months of running an agent-built pipeline unattended, 308 lines of recorded failures next to 208 lines of instructions.
- The Golem Made of English and the Horizon of ConsequencesA language model is a golem made of English. It behaves well only where it can see what its act will cost, and infrastructure is the craft of bringing that cost into view.
Tickets 5
Issues on the board, open first, as of the last build.
Other topics
Built by build_hubs.py from site/topics.yml, the docs labels, the tag pages on the map and the board. Machine-readable: topics.json.