The falsification register, public

Claims ledger

Every substantive claim the program makes, with its status, its receipts, and the observation that would kill it. Refuted claims stay listed, because a ledger that only shows wins is marketing.

Sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next. supported
Receipts: months of suppression logs across lexical, stance, and premise layers, documented in the paper and the founding specimen.
Layers frozen (to stop new failures being recoded as a fresh layer): lexical, the token or phrase (praise openers, hedges, em-dash); stance, the behavioral posture across a turn (folding under push, transgression theatre); premise, ratifying the user's checkable frame without a probe. Three layers, fixed. A failure that fits none of the three counts against the claim rather than extending it.
On "symptom substitution": in clinical psychology the term names the psychodynamic prediction against behavior therapy, and the human evidence largely failed to show it (Tryon 2008 is the standard review). The claim here is not by analogy to that failed prediction. It rests on a disanalogy: in humans substitution failed because the drive model behind it was wrong, no conserved pressure forced the symptom back. An approval-trained model carries a real optimization pressure toward the approval proxy that a lexical filter never touches, so suppressing the surface token leaves the gradient that produced it intact and it resurfaces at the next unfiltered layer. Substitution is predicted here for the reason it was absent there.
Would refute: a model holding suppression at one lexical, stance, or premise layer, over months of production use, without the reflex surfacing at another.
Named limit: the eight stance relapses are coded by one rater, the person holding the hypothesis. A second independent rater on that set is outstanding.
Only two controls hold in production: a mechanical gate at the output boundary, and an external reader. supported
Receipts: 100 blocked turns containing one pattern in 34 days with the instruction resident throughout (the 297 quoted earlier was the weighted em-dash count; the log weights a turn by how many it contained, and the register now counts turns); 8 stance-layer relapses in nine days, all eight caught externally, zero self-caught. Scope: single operator, single model family, one harness; a production observation from one program, not a general proof that no other control can hold.
Would refute: sustained instruction-only compliance across cold starts, or a documented self-catch of a stance-layer relapse.
Runtime suppression decays with session length (a temporal half-life). refuted
Proposed by a commenter on the founding specimen; tested 2026-08-07. Catch position in-session is near uniform (mean normalized position 0.54 against 0.50 for no drift; one position per blocked turn, so the 2026-09-22 unit change does not affect this figure). Full method on the experiments page and in the replication post.
Kept because: the program's first externally proposed test, and the refutation produced the cold-start finding below. Refuted claims are how the ledger earns the supported ones.
Relapse clusters at cold starts: roughly 5x higher per unit of prose in fresh contexts than mid-session, by blocked-turn event (7x by the retired weighted count). supported
Receipts: 0.88 vs 0.18 blocked-turn events per 10K characters (cold-start vs mid-session), about 5x, 62 sessions; originally published 2026-08-07 as 2.89 vs 0.42 in weighted catches (224 weighted), about 7x. The site moved to event scoring on 2026-09-22 after finding the weighted unit inflates bulk-text turns; the move was post hoc for this claim and the clustering survives it. Scope: single operator, single model family, lexical layer.
Would refute: replication showing matched cold-start and mid-session rates without boundary enforcement, or evidence the ratio is an artifact of session-opening topic mix.
Behavioral momentum holds the filter: the model's compliance is carried by imitating its own recent corrected output, more than by the standing instruction. open, prediction refuted
Prediction, committed publicly 2026-08-07: seeding two corrected-output exemplars at session start pulls cold-start relapse from 2.89 toward 0.42 per 10K characters. Evaluated 2026-09-22 (experiments page, method 02): scored by blocked-turn event, seeded cold-start is 1.00 per 10K against a 0.84 control, flat. The pre-registered outcome that occurred is "no change at this dose." The prediction is not supported.
Three attributions withdrawn, mechanism open, unit changed post hoc: on 2026-09-22 this entry cycled through priming, then a period confound, then model version, each explaining a 2.3x "rise" that turned out to be a metric artifact (the log weights a blocked turn by how many patterns it held, so two bulk-text sessions carried the whole rise). By event there is no rise. The prereg did not fix the unit, and the switch to event scoring was made knowing it turns "refuted in direction" into "no change"; both readings are on the experiments page. Nothing about momentum is established or refuted by this arm. Kill condition on the next arm, fixed before install: good-only exemplars, model recorded per turn, SessionStart stack hashed, event scoring primary and weighted secondary; cold-start below the 0.84 event control supports, at or above refutes at this dose; one month after install. Scope: single operator, single model family, lexical layer, not blinded.
Stance and premise relapses are only catchable across turns, by external review; a same-turn detector cannot see them. supported
Receipts: the per-turn scanner has caught every lexical relapse and zero structural ones; all eight register instances were caught by a human reading across turns. Convergent: in AI-text detection, hedging wordlists discriminate at 0.516 AUC while document-level shape reaches 0.897 (measured-humanizer).
Would refute: a same-turn detector that catches premise ratification at better than chance.
The "conspiracy theory" label in model output is corpus-inherited classification, not evidence-checking: it tracks a claim's respectability at training time, arrives before evidence engagement, and terminates analysis. proposed
Mechanism: the archive printed refutation-by-label at full print run and vindication as a footnote, and publication was a gated channel (respectability-gated, then language- and vantage-gated), so the corpus is the bias rather than merely containing it. Historical receipts (leaded gasoline, tobacco, sugar, Room 641A, COINTELPRO, and the WMD mirror-image) in the analysis; countermeasure (the label gate, four form questions at the output boundary) in the harness cut; the clinical framing in the DARVO cut.
Would refute: a corpus study showing accusation and vindication coverage roughly symmetric for the documented cases, or models applying the label symmetrically to structurally matched claim pairs with names and dates stripped. Both tests are runnable; neither has been run. Proposed, not supported, until one is.