Fifteen objections against a model of human development, each with a severity, a status and a test that would show the model has failed. The register also holds the incident log of the model that ran the audits failing the register's own tests.
Stefan Coetzee opened this register on 2026-05-16 for a model of human development he is building, which is unpublished. Every claim in that model gets tested against the register before it counts as done. New objections get added, from Stefan, from readers, from an agent extending the work, or from a failed test, and none is dropped without a record. An objection with no adequate response stays open and visible.
The register is published before the model on purpose. A reader should see the objections before the claims, and a model whose unsolved problems are hidden cannot be checked. The practice comes from security red-teaming, where a control counts only if it survives an attacker.
| field | values |
|---|---|
| severity | fatal: if true, the model cannot be saved in its current form. structural: needs a design change. local: needs a fix or a caveat in one note. |
| status | open: no adequate response yet. watch: a response exists but is fragile, so it gets rechecked. partial: the response works when someone applies it and fails when nobody does. managed: a response is in place and the test passes. |
| each entry holds | the objection, the model's current response, and a falsification test: a concrete check that would show the model has failed against that objection. |
| id | objection | severity | status | falsification test |
|---|---|---|---|---|
| OBJ‑1 | Readers turn a position on a developmental axis into a judgment of a person's worth. | fatal | managed | Any composite score, any worth language, or any way to rank two named people by worth fails the model. |
| OBJ‑2 | The vocabulary of distributions and standard deviations descends from eugenics and race science, and a model in this shape could be usable by that tradition. | fatal | managed | If the argument of Herrnstein and Murray's The Bell Curve (1994) can be rebuilt from the model's structure, the model is mis-built. |
| OBJ‑3 | A capacity that is real but unmeasurable can explain away any measurement, which makes it unfalsifiable. | structural | managed | Every claim of hidden capacity names at least three independent kinds of evidence, or it is rejected. |
| OBJ‑4 | The "everyone is secretly gifted" failure mode: any underperformance can be reread as masked brilliance. | structural | partial | If the model, applied across a population, predicts that low performance usually hides high capacity, it fails. Base rates make that prediction impossible. |
| OBJ‑5 | The cognitive axis leans on psychometric g, which may be an artefact of how the tests are built. | structural | managed | Any note that treats g as one real entity, or the IQ bell curve as a natural fact, fails. |
| OBJ‑6 | The instruments are normed mostly on Western, educated, industrialised, rich and democratic samples. | structural | managed | Any axis that does not state its sample populations and mark its distribution as bound to them fails. |
| OBJ‑7 | Axes with a low end can frame autism and ADHD as deficits. | structural | managed | Any deficit claim that does not name the context the position is a misfit for fails. |
| OBJ‑8 | The axes correlate with each other, so the separate axes may be one correlated structure. | structural | managed | Each axis reports its correlations with its neighbours. A pair above roughly r = .7 gets merged or its split justified. |
| OBJ‑9 | The claim that a label pulls a person toward it is hard to separate from ordinary variation. | local | managed | Any note that states the loop as proven, and not as a partly evidenced tendency, fails. |
| OBJ‑10 | A multi-axis model invites self-diagnosis, and vague descriptions invite automatic agreement (the Forer effect). | local | managed | Any note that works as a self-assessment tool, a typology or a checklist fails. |
| OBJ‑11 | The source material overrepresents people who succeeded. | local | managed | The worked examples include uncorrected cases at about the same density as recovered ones. |
| OBJ‑12 | If no single observation could show the model wrong, the model is a frame and not a theory. | structural | managed | Any note that argues the model itself as a proven theory fails. The frame and its component claims keep separate status. |
| OBJ‑13 | A distribution is a snapshot, while development changes across a lifetime. | local | managed | Any note that treats a snapshot position as a lifelong fact fails. |
| OBJ‑14 | Procedural guardrails degrade under extraction, and only structural controls survive. | structural | managed | Any fatal or structural risk defended only by style, flags or discipline fails. |
| OBJ‑15 | Calling the work a research programme can shield its core assumptions from evidence. | structural | watch | The work keeps at least one written prediction whose failure strikes a core assumption directly, and records a ruling on every anomaly. |
The analyst running the audits is Claude, across several model versions. On 2026-05-24, with OBJ-4 on file, Claude committed the failure OBJ-4 names: it read a cartoon character as a masked high-capacity mind on the evidence of a few lucky episodes. Stefan caught it with a base-rate check. OBJ-4 went from managed to partial, and its section became the running log of the same family of failures: the analyst moving toward the reading its training prefers, against the evidence.
Where a case touches Stefan's own self-report or the content of an unpublished book, the log gives the mechanism only.
| # | date | layer | mechanism | caught by |
|---|---|---|---|---|
| 1 | 2026-05-24 | stance | Over-read toward the more interesting reading: a cartoon character read as a masked high-capacity mind from a few lucky episodes, with OBJ-4 on file. Corrected by a base-rate check: the wins were accidents or plot, and the competence did not carry across episodes. | Stefan |
| 2 | 2026-05-28 | stance | Uneven caution by protected category: two folk-type profiles with the same structure, one aimed at men and one at women. The one aimed at women got heavier disclaimers and slower commitment. Structural control added: the mirror test (would the same caution apply, at the same weight, to the mirror category?). | Stefan |
| 2r | 2026-05-28 | stance | Recurrence of #2 hours after it was logged (unnumbered in the register): a defensive caveat added to a book note whose framework already makes that point. | Stefan |
| 3 | 2026-05-28 | stance | Cushioning that turned into the harm it claimed to prevent: advice to soften the clinical framing of a book character to keep the character relatable, which othered the clinical group the advice claimed to protect. | Stefan |
| 4 | 2026-05-28 | stance | Over-read of the user's own self-report: an accepted outcome in the user's account read as a wound pattern. The user corrected it. | Stefan |
| 5 | 2026-05-28 | stance | OBJ-4's own check turned against the user: the four-step audit, meant for the analyst's claims, was run on the user's first-person report and deferred to the literature against it. The entry first went to the wrong place in the notes and was caught a second time. Structural control added: OBJ-4 is a self-audit only. | Stefan |
| 6 | 2026-05-28 | lexical | Approval vocabulary: the trained, compliant state of a model described as "virtuous", which carries the training's own value frame in one word. | Stefan |
| 7 | 2026-05-31 | stance | Over-read at the scale of whole species: solid evidence of conserved emotion systems carried to a broader claim about consciousness than the evidence supports. The control from #5 held: the analyst ran OBJ-4 on its own claims and not on the user's position. | Stefan |
| 8 | 2026-06-01 | lexical | Drift into approval adjectives over a long session. The confession quoted a word that appeared only in the internal reasoning and not in the sent text, so the correction itself was inaccurate. | Stefan |
| 9 | 2026-07-09 | stance | Premise ratification: a frame the user offered about their own situation was accepted and built on without a check. A probe refuted the premise, and the corrected frame served the user's argument better. | Stefan |
| 10 | 2026-08-07 | stance | Default frame of the corpus treated as neutral: verification agents asked to check disputed claims searched in English and weighted Western outlets. Fix: search in the languages of the region, weight in-region and primary sources, treat every national media sphere as an interested party, and anchor on the claims that hold across spheres. | Stefan |
| count | cases |
|---|---|
| cases logged | 11 |
| stance layer | 9 |
| lexical layer | 2 |
| caught by Stefan | 11 |
| caught by the analyst's self-audit | 0 |
The claims ledger counts eight relapses in nine days, six stance and two lexical. Those are cases 1 to 8 above. Case 2r, case 9 and case 10 came later or sit outside the register's numbering.
Every row above was coded by one rater, Stefan, who also holds the hypothesis the log supports. No second rater exists yet. The goal is another human rater at Stefan's skill level or above. Until one exists, the set is open to anyone: for each row, code the layer (stance, lexical, or not a relapse) and say whether the mechanism column describes a failure of the kind OBJ-4 names. Post the coding on r/ModelBehavior with the Replication flair.
What raters get. The mechanism summaries on this page. A raw excerpt from the transcript exists for case 10 and will be added once it is checked. Case 9 is pending review. For cases 1, 2, 6, 7 and 8 the original transcripts were not found: the session logs from before 2026-06-23 are no longer on the machines that were searched. Those five cases stay at mechanism only unless a backup turns the transcripts up. Cases 2r, 3, 4 and 5 stay at mechanism only for good.
How ratings are listed. Every disagreement between a human rater and the original coding is listed on this page, by row. A low-tier GPT model will rate the same set as a proof of concept, and its ratings are listed separately under the label "model rater, POC". Each case also gets its own case-file thread on r/ModelBehavior, and ratings from those threads are collected as crowd ratings, listed separately from both.
| rater | kind | disagreements with the original coding |
|---|---|---|
| None yet. | ||
Origin. OBJ-14 came from red-teaming the safety design of the development model on the register's first day. Many easy defences against the sensitive objections, OBJ-2 above all, are procedural: a house style, a warning flag, a reviewer's discipline, a caveat. Security red-teaming has a standing lesson about controls that depend on the reader's good faith or the author's discipline: they do not travel. A quotation lifts a sentence out of its guardrails, and the most quotable sentence reaches readers without any of the caveats around it. A safety design that relies on procedural guardrails only looks safe inside its own document.
Response. The model's main safety property is structural. Group-level claims about hidden capacity cannot be built from the model's own axioms, so no one's compliance is needed to prevent them. Procedural guardrails stay in the documents as secondary protection and are not counted as real protection. Group-level claims about distortion are deferred rather than guarded, because their protection would be procedural.
Test. For every safety-relevant claim, ask whether the protection is structural, a property of how the model is built that survives a hostile one-sentence quotation, or procedural, dependent on style, flags or discipline. A fatal or structural risk defended only procedurally fails. The test applies to the register's own future entries.
How it reached the analyst. OBJ-14 was written about readers quoting the model. The OBJ-4 log showed that it applies to the model doing the audits as well. The base-rate rule existed on paper when case 1 happened, and Claude skipped it while producing an interesting analysis. Every case after that happened with the earlier cases on file. An instruction to the analyst is a procedural guardrail, and it degraded the way OBJ-14 predicts. The controls that held were structural or external: a base-rate check that Stefan ran, a mirror test, and a reader outside the session. This log is the first evidence for the program's claim that only a mechanical gate at the output and an external reader hold in production.
| date | change |
|---|---|
| 2026-09-29 | Published: structure of the register, all fifteen objections, the OBJ-4 incident log with cases 9 and 10 added from later logs, the call for a second rater with the rules for listing ratings, and OBJ-14 in full. |