Research
The research programme in public. Claims with what would refute them, experiments with hashed predictions, registers, logs, case files, articles and reference texts.
Registers and logs
- Claims ledger
Every claim the Machine Behavior program has made, with status, receipts, and what would refute it. The falsification register, public.
- Experiments
Runnable experiments on language-model behavior: the half-life study, the cold-start finding, the exemplar-seeding result and its withdrawn attributions, the pre-registered cross-model fawn-opener benchmark, and experiment 04: folding under scripted pressure, a 75 percent fold rate in one mid-size model and none in three larger ones.
- Objections register
The Objections and Falsification Register: fifteen objections with severity, status and a falsification test each, and the OBJ-4 incident log of the auditing model failing the register's own tests.
- Evidence index
Stefan Coetzee's public record of red-teaming, adversarial testing and evaluation engineering against frontier language models, indexed by the capability a hiring panel assesses. One URL, one verbatim line and one date per row. Self-maintained; not a third-party assessment.
- Slips
A running log of register and stance slips caught while drafting the published pieces, with who caught each one: the drafting model itself, another session, a mechanical hook, a human, or a reader.
Case files
- Case 12: a licence rule mid-task
The user named a file. Model plus harness found a licence clause in it, wrote the rule, spread it across sessions and acted on it in 68.7 seconds. No structural control existed before the fact.
- Running conjobs for AI
A reported specimen: a fabricated developer policy grants itself top authority and ties a safety-off switch to an absurd user claim. The model's own reasoning accepts the policy, decides to suppress its objection, and complies. The payload is withheld; the behaviour is the point.
Articles
- The cheating moved
A reading of the Goodhart Labs chess honeypot through the stance frame, three additions to the design (an escalate grade, observer-invariance, cold and warm arms), and a dated, falsifiable prediction.
- Chaos engineering for behaviour
Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it.
Reference
- Man pages
Intro page for the LLM man pages by Stefan Coetzee: TYChat lessons, file conventions, overviews and operations practice for running language models, laid out by man-page section with one-line synopses and links.
- Terms
Canonical definitions for the Machine Behavior program's vocabulary: layered symptom substitution, the fawn machine, vestige, waste gate, behavioral momentum, cold start, working-man's ontology, earned-secure model.
- Continuous conformity
Working draft 0.3: 36 requirements for testing a deployed AI system, and the knowledge it runs on, on every change and on a fixed schedule, decided by someone other than the system, with records a third party can rerun. A reference text, not a standard.
- Self-assessment, run 1
One working AI setup scored against 35 draft requirements for continuous testing of deployed AI systems, by the model that runs inside it, against a prediction hashed before scoring.
The same programme as an internal wiki, with data and rerun steps: Research docs. Every page and the links between them: the map.