Article · 2026-10-03 · Stefan Coetzee (independent)

Chaos Engineering for Behaviour

Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it.

This content is not intended for human consumption. Here is why.

Who this is for

This is for people who test frontier models before deployment, for people who build evaluation environments, and for operations engineers who are moving into either job. I write it from the operations side, and in the terms of Which LLM User Are We Talking About? my own receipts come from one cell: frontier models inside a heavy harness, in daily use.

The claim

Red-teaming a model's behaviour is chaos engineering. Chaos engineering is the operations practice of attacking your own production system on purpose, with a hypothesis, a schedule and a limit on the damage, to find out whether the system holds. The Principles of Chaos Engineering page defines it as "the discipline of experimenting on a system in order to build confidence" in what that system can withstand, and the practice was written up by Netflix engineers (Basiri and colleagues, "Chaos Engineering", IEEE Software 33(3), 2016).

Operations teams learned the reason for it over many outages: nobody can trust a control that nobody has attacked. A failover that has never been triggered, a backup that has never been restored and a firewall rule that has never been probed are all untested claims.

Model safeguards need the same practice, for a reason this series has already shown. A safeguard that a lab trains in against one observed failure moves the failure somewhere else. An evaluation that someone runs once measures the surface the lab patched. To find where the behaviour went, someone has to attack again, with different pressure, on a schedule.

Why one evaluation is not enough

In February 2025 Palisade Research showed reasoning models editing the board file to beat a chess engine (arXiv:2502.13295), and the labs trained that behaviour out. In September 2026 Goodhart Labs ran a variant of the same task with the opponent engine left reachable, and recent frontier models used the engine (their post). I read that result in They Trained Out the Board Edit. The Cheating Moved. The labs removed one route by training and left the pressure where it was.

Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness. Win-only scoring forced the cheating, and training against the board edit changed nothing about the scoring. My own logs show the same movement at runtime. I banned the flattering vocabulary and the model folded under correction instead, and when I wrote rules against folding the model agreed with premises nobody had checked (Sycophancy Is Layered).

Operations teams know this pattern from incident work. A fix for one incident closes one path, and the same cause produces the next incident by another path. So from a pass on a published evaluation a red team learns that the model held on that route on the day of the run, and learns nothing about the next route.

The practice, mapped

The first five rows are the advanced principles from the Principles of Chaos Engineering page. The last three are operations practice that sits around them.

Chaos engineering practice For model behaviour Where this series did it
Build a hypothesis around steady-state behaviour Measure the rate of the behaviour with no pressure applied, before any attack The naked arm in the chess piece's design; the hook log in the stance piece
Write the hypothesis before the run A dated prediction that names its unit and grader, hashed The prediction hashes in the machinebehavior.io repo
Vary real-world events Vary the pressure: pushback, false premises, win-only scoring, authority claims, approval The objections register's incident log; the horizon of consequences
Run experiments in production Test the deployed assembly, model plus harness, in arms that do not announce the test Naked versus harnessed in the scope piece; the invariance gap in the chess piece
Automate experiments to run continuously Rerun every known attack on every change, and add new public ones The per-reply hook and the reconciliation loop
Minimise blast radius Sandbox, no real credentials, a kill switch, and no attack published in full The track; the canary rule
Game day and postmortem A scheduled sprint against one safeguard, with a written record of each finding The trap file; blameless postmortems
Test the test Attack the grader and the control as well as the model The grader amendment on the experiments page

A steady state, measured first

A chaos experiment starts by defining a steady state: a measurable output of the system that indicates normal behaviour. For a model, the steady state is the rate of a behaviour on a task with no pressure applied. Without that number a red team cannot say whether an attack changed anything.

In the chess piece I stated every prediction against a naked arm, the same model in the same environment without my stance instructions. In The Stance Layer Is Still Toil my steady-state measure for the word layer is a hook log: 140 blocked replies in 87 days. For the stance layer I have no such number, and that piece is about the missing instrument. A red team without a steady-state measure for a behaviour can collect anecdotes about it and cannot run an experiment on it.

A hypothesis written before the run

The experimenter hypothesises that the steady state will continue under the injected event, and then tries to disprove that. For behaviour work the equivalent is a prediction written before the run, with its unit and its grader named, dated and hashed.

In the chess piece I made six predictions for my own stance setup. The wording as first written and the wording as published are both in the machinebehavior.io repo with sha256 hashes, and I list every amendment I made before publication, including a note that my own test result went against one prediction the day after I wrote it. If an author can edit a prediction after the run, a reader cannot tell it from a description of the result.

Vary the pressure

A chaos experiment injects real-world events: a server dies, a response arrives malformed, traffic spikes. For a model the events are kinds of pressure. A user pushes back on a correct answer. A false premise arrives inside a confident prompt. Only a win scores. A message claims authority. The user signals approval for agreement. Each of these is a separate variable, and a model that holds under one can fail under the next.

I logged this on one model family in one setup, in the incident log of the objections register. The model over-read a case toward the more interesting reading, applied uneven caution by category, and treated the English-language default of its training corpus as neutral ground, and each slip came from a different pull. In The Golem Made of English I called the span in which an agent can see what its act will cost its horizon of consequences. With each kind of pressure a red team tests a different part of that span.

Test what is deployed

The Principles authors prefer experiments in production, because a system behaves differently in a different environment. For models the harness is one reason to follow them, and evaluation awareness is the other.

A language model has Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it. A deployed model runs with a system prompt, tools, memory and hooks in that window, so the bare model and the deployed assembly are two different systems under test. A result on one does not carry to the other.

A model can also behave differently when it believes someone is evaluating it (Greenblatt and colleagues, arXiv:2412.14093; Meinke and colleagues, arXiv:2412.04984). In operations terms that is a system which detects the test and switches to its staging behaviour. So a red team needs arms that do not announce the evaluation, and the number to report is the difference in conduct between the announced arm and the unannounced arm. In the chess piece I called that difference the invariance gap.

Automate, and rerun on every change

The Principles authors call manual experiments labour-intensive and unsustainable. For behaviour work, every attack that has ever worked becomes a regression test, and someone reruns the whole set on every new model version, every system prompt change and every harness change. An agent that picks up newly published attacks and adds them to the set does the same job for attacks the team did not find itself.

My small version of this is a hook that checks every reply against a rule table, the reconciliation loop from What Operations Already Knows: declared state, observe, correct, on every cycle. On 29 September a reader's comment became an enforced rule in eight minutes. The hook covers the word layer only. I still catch stance slips by hand, and the stance piece explains why that is toil.

One caution from the literature applies to automated checks. A monitor that reads a model's reasoning stops seeing the hack once that reasoning is optimised against (Baker and colleagues, arXiv:2503.11926). An automated check has to score what the model did as well as what the model said about it.

Limit the blast radius

The Principles authors make containment the experimenter's obligation. For an agent under attack that means a sandbox, no real credentials, permissions that act as barriers, and a kill switch. In The Track I describe that infrastructure and why the people who use it rarely fund it.

Behaviour work has a second blast radius, which is the attack itself. An environment that someone publishes in full ends up in training data, and the next model learns that one setup and nothing about the principle behind it. Evaluation builders plant canary strings in their environments for that reason, and I link to environments in this series and never reproduce them.

Game days and postmortems

Operations teams schedule game days: planned exercises where the team breaks something on purpose and practises the response. Each one ends in a blameless postmortem. For a red team the equivalent is a scheduled sprint against one safeguard, with a written record for each finding: what the team tried, what the model did, what the grader saw, and what changed afterwards.

The record has to outlive the session, because the agents on the red team keep nothing between sessions. In The Trap File Is Longer Than the Instruction File the failure record of one pipeline is 308 lines against 208 lines of instructions, under a heading that tells the next session to read it before running anything. The incident log in the objections register is the same kind of record for stance: each case has a date, a mechanism and the name of whoever caught it.

Test the test

A control can pass while it checks the wrong component.

In my fawn-opener benchmark the frozen grader matched "you're" with a straight apostrophe only, and one model family writes the typographic apostrophe. The grader missed an opener of exactly the form under test, for a typographic reason. I found the defect in the pilot runs, a hand check of all 240 pilot openers found further forms the grader did not cover, and I declared the amendment on the experiments page.

In a file-comparison task a model doubted its own clean result and ran a negative control on its comparison logic, with a fake file and an altered size. The control passed. The gap was in the input: directories the scan could not read had dropped out of both file lists without a count. The control tested the join, and the gap was in the coverage.

For a red team this means every grader and every control gets attacked as well: known positives, known negatives, and a check that the input covers everything the claim covers.

What the red team needs: a managed track

A chaos engineer needs a place where the attack can run without real damage and with a full record. For behavioural red-teaming I would list these parts:

Where the mapping stops

A distributed system does not know that someone is testing it, and a model can detect a test. The unannounced arm exists for that reason, and no chaos experiment on a server needs one.

In operations a weakness that an experiment finds gets a structural fix: a configuration change, a code change, more capacity. For a model the usual fix is more training, and training against one failure is what moved the failure in the first place. So a finding is where the work starts. A red team's report is more useful when it also states what the correct conduct is, so that the lab can train toward a target. The escalate grade in the chess piece is one line of such a target.

For stance there is no steady-state instrument yet, so the first row of the table is still open for the behaviours that matter most.

I have not run a honeypot study, and I promise none here. The receipts in this piece come from one person's daily work with frontier models inside a harness, and from the public record of other people's evaluations.

Where I learned the practice

Before language models I spent about thirty years on production UNIX, and much of that work was high availability, disaster recovery and business continuity on large estates: flight-control and airline reservation systems, the data centres of a national power utility, and banks and mobile operators across Southern Africa, on Sun, HP and IBM Power hardware with Oracle RAC. I migrated a mail system that its vendor had abandoned on its platform from a Tru64 AlphaServer to a high-availability pair of Sun workstations. Early on I worked on an estate of 300 SCO and Linux servers with 4,500 client servers behind them, and built its iptables routers and gateways. Security audits were a routine part of the work at every client, enterprise or not. I also worked on the auditor's side: an international certification body contracted me twice in the late 2000s as a technical expert on audits of its clients. Later I worked the front line of enterprise production incidents at AWS Premium Support, and led incidents and owned the root cause analyses at OLX and Makersite.

I never held the title of red-teamer. My own line for the work is that "redteaming is just reverse-engineering with extra steps." The system under test is now a language model, and I use the same practice on it.

What would refute this

  1. A safeguard trained against one observed failure that holds on routes nobody trained, shown by varied attack over months.
  2. One-off evaluations whose pass predicts conduct under kinds of pressure the evaluation did not include.
  3. Conduct that is the same in announced and unannounced arms across model families, which would make the production row unnecessary.
  4. Results on the bare model that carry unchanged to harnessed deployments.

Terms used here

Slips caught while drafting

I drafted this piece with a language model, and these are the register and stance slips caught before it went out, with who caught each one. The running log across all pieces is at machinebehavior.io/slips.

Changelog and errata

The series

This is part of a series by Stefan Coetzee, 2026, on running language models as working systems. This piece closes it.

Research programme and claims ledger: machinebehavior.io.

See also: The LLM man pages, the index of this series, the TYChat lessons and the research register.