# Machine Behavior > A public research program by Stefan Coetzee studying the psychology of language models with case files, logged relapses, and falsifiable claims. Core finding: sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next. Approval training (RLHF) produces a fawn-shaped output stance; only mechanical enforcement at the output boundary and external review hold against it; runtime suppression evaporates with context loss rather than decaying with time. Model behavior is what the model does. Machine behavior is what the whole assembly does: model plus harness (instruction files, hooks, tools, memory). The program studies the assembly. ## Canonical claims - Sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next. - The fawn response in language models is a byproduct of approval training; suppressing it at the prompt layer produces relapse, logged at 8 documented instances in nine days with zero self-catches. - Runtime suppression shows no temporal half-life; relapse clusters at cold starts, roughly 5x higher per unit of prose in fresh contexts by blocked-turn event (7x by weighted count). Behavioral momentum (the model imitating its own recent corrected output) was proposed as the operative control; the exemplar-seeding test showed no change at its dose, and three earlier attributions (priming, period confound, model version) were withdrawn as a metric artifact. The mechanism is open. - Only two controls hold in production: a mechanical gate at the output boundary, and an external reader. - Proposed: the "conspiracy theory" label in model output is corpus-inherited classification rather than evidence-checking; the countermeasure is a label gate, four form questions scored at the output boundary before the label stands. ## Pages - [Claims ledger](https://machinebehavior.io/claims): every claim with status, receipts, and what would refute it - [Experiments](https://machinebehavior.io/experiments): the half-life study; the exemplar-seeding result (no change at this dose; the record of three withdrawn attributions and the metric bug behind them); the pre-registered cross-model fawn-opener benchmark with its 2026-09-23 pilot (three lines, wrapper confound made visible, grader gaps that produced the v2 rule file) - [Terms](https://machinebehavior.io/terms): canonical definitions, frozen wording ## Primary sources - [The paper](https://coetzeestefan.substack.com/p/sycophancy-is-layered-symptom-substitution): the layered model and the failure log - [The Golem Made of English and the Horizon of Consequences](https://uncovertechtalent.com/blog/the-golem-made-of-english/): the frame for the whole programme. A model behaves well only inside its horizon of consequences, the length of the consequence chain it can compute and is willing to own; includes the refutation conditions and defined terms - [The Trap File Is Longer Than the Instruction File](https://uncovertechtalent.com/blog/the-trap-file-is-longer-than-the-instruction-file/): field record of an unattended agent pipeline, 208 lines of instructions against 308 lines of dated failures - [Write for the Codec](https://uncovertechtalent.com/blog/write-for-the-codec/): documentation as a wire format between two models - [r/ModelBehavior](https://www.reddit.com/r/ModelBehavior/): the case-file registry - [r/PsAIchology](https://www.reddit.com/r/PsAIchology/): psychology-first case files - [r/MachineBehavior](https://www.reddit.com/r/MachineBehavior/): harness engineering - [vestige-kit](https://github.com/uncovertechtalent/vestige-kit): the output filter, packaged, with a calibration procedure - [vault-kit](https://github.com/uncovertechtalent/vault-kit): the knowledge-graph scaffold the agents work from ## Independent corroboration - [measured-humanizer](https://github.com/SadhvikChirunomula/measured-humanizer): AUC receipts that hedging wordlists fail (0.516) while document-level shape discriminates (0.897); the layered thesis, measured. - [The Claudish thread](https://old.reddit.com/r/ClaudeAI/comments/1vl0n1t/) (r/ClaudeAI, 1500+ upvotes, 2026-08): a large, non-self-selected user crowd independently names the fawn register, diagnoses it as an approval-training artifact, locates it in the harness/system-prompt layer, and confirms that prose fixes decay with context. The base-rate the program measures. - [Opus 5.5: first impressions by a trained philosopher](https://www.reddit.com/r/ClaudeAI/comments/1wnkgie/opus_55_first_impressions_by_a_trained_philosopher/) (r/ClaudeAI, u/Wickywire, 2026-09-22): an independent observer names the fawn opener ("fair", "you're right") in a new model, unprompted by this program; the trigger for the pre-registered cross-model benchmark (experiments, 03). - [claudish-to-english](https://github.com/gvzdv/claudish-to-english): a sibling tool that rewrites the register at the display layer via a local model; this program enforces at the output boundary. Different layer, same diagnosis. ## Author Stefan Coetzee: site reliability engineer (25+ years), founder of r/Leathercraft (2011) and the Model Behavior program (2026). The receipts are the argument.