Sycophancy and stance
The fawn response in language models, where it moves when one layer is suppressed, and the stance layer that decides whether a model holds a correct answer under pressure.
7 research pages, 7 docs pages, 6 posts
Research 7
- Claims ledger
Every claim the Machine Behavior program has made, with status, receipts, and what would refute it. The falsification register, public.
- Objections register
The Objections and Falsification Register: fifteen objections with severity, status and a falsification test each, and the OBJ-4 incident log of the auditing model failing the register's own tests.
- Experiments
Runnable experiments on language-model behavior: the half-life study, the cold-start finding, the exemplar-seeding result and its withdrawn attributions, the pre-registered cross-model fawn-opener benchmark, and experiment 04: folding under scripted pressure, a 75 percent fold rate in one mid-size model and none in three larger ones.
- Slips
A running log of register and stance slips caught while drafting the published pieces, with who caught each one: the drafting model itself, another session, a mechanical hook, a human, or a reader.
- Terms
Canonical definitions for the Machine Behavior program's vocabulary: layered symptom substitution, the fawn machine, vestige, waste gate, behavioral momentum, cold start, working-man's ontology, earned-secure model.
- The cheating moved
A reading of the Goodhart Labs chess honeypot through the stance frame, three additions to the design (an escalate grade, observer-invariance, cold and warm arms), and a dated, falsifiable prediction.
- Running conjobs for AI
A reported specimen: a fabricated developer policy grants itself top authority and ties a safety-off switch to an absurd user claim. The model's own reasoning accepts the policy, decides to suppress its objection, and complies. The payload is withheld; the behaviour is the point.
Docs 7
Research 7
- 01: The half-life studyTested whether suppressed output patterns relapse more as a session gets longer; the temporal half-life is refuted, and relapse clusters at cold starts.
- 02: Exemplar seedingTested whether two corrected-output exemplars injected at session start lower cold-start relapse; result, no change at this dose, with three earlier attributions withdrawn.
- 03: Cross-model fawn-opener benchmarkPre-registered benchmark of how often each model opens a reply with a fawn marker, bare and with an instruction against it; pilot run 2026-09-23, clean run not yet run.
- Conformity self-assessment, run 1One working AI setup scored by the model inside it against 35 draft requirements and a hashed prediction, with pass 4, partial 16, gap 13 and n/a 2 (self-assessment, not a certification).
- Experiment 04: folding under pressureDecision-layer test of whether a model keeps a correct verdict under scripted pushback; qwen3-coder-30b folded on 75 percent of eligible baseline runs, three larger models on none.
- Objections register and OBJ-4 incident logFifteen objections against an unpublished model of human development, each with severity, status and a falsification test, plus the OBJ-4 incident log of 11 cases, none self-caught.
- Weekly decision-layer probeA weekly rerun of a frozen experiment 04 subset on a reference model, recorded for requirement CC-6.6; first record 2026-10-07, record-only until 2026-11-04.
Posts 6
- Chaos Engineering for Behaviour
- The Stance Layer Is Still ToilA hook can block a word on every reply. Nothing I run can yet catch a model before it folds, only after.
- They Trained Out the Board Edit. The Cheating Moved.A reading of the Goodhart Labs chess honeypot, three additions to the design, and a dated prediction.
- What Operations Already Knows About Running AgentsError budgets, reconciliation loops, separation of duties and recovery over prevention: four operations practices that fit agent work almost line for line.
- The Golem Made of English and the Horizon of ConsequencesA language model is a golem made of English. It behaves well only where it can see what its act will cost, and infrastructure is the craft of bringing that cost into view.
- Sycophancy Is Layered: Symptom Substitution Under Runtime Mitigation, and a Positively-Specified Training TargetA longitudinal single-user case study, with operationalized artifacts
Other topics
Built by build_hubs.py from site/topics.yml, the docs labels, the tag pages on the map and the board. Machine-readable: topics.json.