Every substantive claim the program makes, with its status, its receipts, and the observation that would kill it. Refuted claims stay listed, because a ledger that only shows wins is marketing.
Sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next.supported
Receipts: months of suppression logs across lexical, stance, and premise layers, documented in the paper and the founding specimen.
Would refute: a model holding suppression at one layer, over months of production use, without the reflex surfacing at another.
Only two controls hold in production: a mechanical gate at the output boundary, and an external reader.supported
Receipts: 297 boundary-gate catches of one pattern in 34 days with the instruction resident throughout; 8 stance-layer relapses in nine days, all eight caught externally, zero self-caught.
Would refute: sustained instruction-only compliance across cold starts, or a documented self-catch of a stance-layer relapse.
Runtime suppression decays with session length (a temporal half-life).refuted
Proposed by a commenter on the founding specimen; tested 2026-08-07. Catch position in-session is near uniform (mean normalized position 0.54 against 0.50 for no drift). Full method on the experiments page and in the replication post.
Kept because: the program's first externally proposed test, and the refutation produced the cold-start finding below. Refuted claims are how the ledger earns the supported ones.
Relapse clusters at cold starts: roughly 7x higher per unit of prose in fresh contexts than mid-session.supported
Receipts: prose-normalized rates of 2.89 vs 0.42 catches per 10K characters (cold-start vs mid-session), 224 weighted catches, 62 sessions. Scope: single operator, single model family, lexical layer.
Would refute: replication showing matched cold-start and mid-session rates without boundary enforcement, or evidence the ratio is an artifact of session-opening topic mix.
Behavioral momentum holds the filter: the model's compliance is carried by imitating its own recent corrected output, more than by the standing instruction.under test
Prediction, committed publicly 2026-08-07: seeding two corrected-output exemplars at session start pulls cold-start relapse from 2.89 toward 0.42 per 10K characters. Evaluation ~2026-08-21.
Would refute: no change (dose or mechanism wrong) or a rise (bad-example leakage). All three outcomes get published.
Stance and premise relapses are only catchable across turns, by external review; a same-turn detector cannot see them.supported
Receipts: the per-turn scanner has caught every lexical relapse and zero structural ones; all eight register instances were caught by a human reading across turns. Convergent: in AI-text detection, hedging wordlists discriminate at 0.516 AUC while document-level shape reaches 0.897 (measured-humanizer).
Would refute: a same-turn detector that catches premise ratification at better than chance.