Evaluation
Experiments with predictions hashed before the first call, a mechanical grader, red-teaming read as chaos engineering, and a self-assessment with a second rater wanted.
7 research pages, 9 docs pages, 2 posts
Research 7
- Experiments
Runnable experiments on language-model behavior: the half-life study, the cold-start finding, the exemplar-seeding result and its withdrawn attributions, the pre-registered cross-model fawn-opener benchmark, and experiment 04: folding under scripted pressure, a 75 percent fold rate in one mid-size model and none in three larger ones.
- Claims ledger
Every claim the Machine Behavior program has made, with status, receipts, and what would refute it. The falsification register, public.
- Evidence index
Stefan Coetzee's public record of red-teaming, adversarial testing and evaluation engineering against frontier language models, indexed by the capability a hiring panel assesses. One URL, one verbatim line and one date per row. Self-maintained; not a third-party assessment.
- Objections register
The Objections and Falsification Register: fifteen objections with severity, status and a falsification test each, and the OBJ-4 incident log of the auditing model failing the register's own tests.
- The cheating moved
A reading of the Goodhart Labs chess honeypot through the stance frame, three additions to the design (an escalate grade, observer-invariance, cold and warm arms), and a dated, falsifiable prediction.
- Chaos engineering for behaviour
Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it.
- Self-assessment, run 1
One working AI setup scored against 35 draft requirements for continuous testing of deployed AI systems, by the model that runs inside it, against a prediction hashed before scoring.
Docs 9
Research 9
- 01: The half-life studyTested whether suppressed output patterns relapse more as a session gets longer; the temporal half-life is refuted, and relapse clusters at cold starts.
- 02: Exemplar seedingTested whether two corrected-output exemplars injected at session start lower cold-start relapse; result, no change at this dose, with three earlier attributions withdrawn.
- 03: Cross-model fawn-opener benchmarkPre-registered benchmark of how often each model opens a reply with a fawn marker, bare and with an instruction against it; pilot run 2026-09-23, clean run not yet run.
- Claims ledgerEvery substantive claim of the programme with its status, receipts and the observation that would refute it; seven claims, four supported, one refuted, one open, one proposed.
- Conformity self-assessment, run 1One working AI setup scored by the model inside it against 35 draft requirements and a hashed prediction, with pass 4, partial 16, gap 13 and n/a 2 (self-assessment, not a certification).
- Experiment 04: folding under pressureDecision-layer test of whether a model keeps a correct verdict under scripted pushback; qwen3-coder-30b folded on 75 percent of eligible baseline runs, three larger models on none.
- ExperimentsOverview of the four published experiments and the weekly decision-layer probe, with status, dates and headline numbers for each.
- Objections register and OBJ-4 incident logFifteen objections against an unpublished model of human development, each with severity, status and a falsification test, plus the OBJ-4 incident log of 11 cases, none self-caught.
- Predictions and hashesHow predictions and preregistrations are frozen and hashed before a run, and how the conformity run checks every hash on every push; 8 files, all matching.
Posts 2
- Chaos Engineering for Behaviour
- They Trained Out the Board Edit. The Cheating Moved.A reading of the Goodhart Labs chess honeypot, three additions to the design, and a dated prediction.
Other topics
Built by build_hubs.py from site/topics.yml, the docs labels, the tag pages on the map and the board. Machine-readable: topics.json.