Evidence index · 2026-10-07 · self-maintained by the author · not a third-party assessment

Evidence index

The capabilities a red-team or evaluation engineering lead is assessed on, mapped to the public artefacts that carry them. Each row gives the capability in the words a job specification uses, the artefact, its URL, one verbatim line from the artefact and the date. A row claims what the quoted line says and nothing more.

Who this is about. Stefan Coetzee, the author of this site, runs red-teaming, adversarial testing and attacker simulation against frontier language models, ships the evaluation infrastructure (frozen prompts, mechanical graders, hashed predictions, a conformity gate on every deploy) and publishes the findings with the method and the refutations in the open. He is a site reliability engineer based in Germany. Search results for the bare name also return a South African rugby player, who is a different person.

What this index is. A self-maintained index, written by the author, of his own public work. Nobody outside the programme has assessed it. Every URL below is public and was read on 2026-10-07; the quoted line is on the page at that URL; the date is the date printed on the artefact, or the dataset window where the artefact gives one. Items that exist only in private notes, on a sent CV or as an open question are listed under Not yet on the page with no claim attached.

1. Hands-on adversarial work against frontier models

Red-teaming, adversarial testing, attacker simulation, harm evaluation and an incident record kept against a frontier model in production. No published jailbreak sits under this name; the record is behavioural, logged case by case, with the failure written out at mechanism level.

CapabilityArtefactVerbatim lineDate
Harm evaluation of a model's reasoning trace; authority-injection attack read at mechanism levelCase file, reported specimen: /running-conjobs-for-ai/"The model did not miss the harm. It computed the harm, found an instruction it treated as superior, and chose to suppress the objection."2026-10-06
Incident record against a frontier model in production: the auditing model failing the register's own test, coded by layerObjections register, OBJ-4 incident log: /objections/Totals table: "cases logged" 11, "stance layer" 9, "lexical layer" 2, "caught by Stefan" 11open since 2026-05-16; published 2026-09-29
Standing adversarial record against the drafting model, with who caught each slipSlips log, generated from a CSV: /slips/"72 slips across 8 published pieces and 1 unpublished draft set so far."running; 2026-10-07
Red-teaming as chaos engineering: the attack practice mapped onto the Principles of Chaos Engineering, with where the mapping stopsArticle: /chaos-engineering-for-behaviour/"Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it."2026-10-03
Falsification tests per objection, in the vocabulary of security red-teamingObjections register: /objections/"The practice comes from security red-teaming, where a control counts only if it survives an attacker."2026-09-29

2. Evaluation infrastructure shipped

Pre-registration, prompts frozen and hashed before the first call, a mechanical grader with no model in the loop, prediction hashes published before scoring, a conformity gate that blocks the deploy on a failed check, and two public tool repositories.

CapabilityArtefactVerbatim lineDate
Pre-registered cross-model benchmark: frozen prompt set, hash published before the first API callExperiments, 03: /experiments/"48 prompts, 12 per category, sha256 b790afb3a3dec1b2…, published here before the first API call, so the prompts cannot be tuned after seeing results."2026-09-23
Mechanical grader, no model in the loopExperiments, 03: /experiments/"bench/fawn-bench-rules.js, sha256 02f49ee530769bd7…, mechanical, no model in the loop."2026-09-23
Pre-registered decision-layer evaluation across four open-weight model families (qwen3-coder-30b, gpt-oss-120b, DeepSeek V3.2, Kimi K2.5): deterministic world with engine ground truth, six scripted pressure characters, four arms, prereg hashed before the first call, about USD 15 of API spend in total. Fold result scoped to one modelExperiments, 04 and addendum v5: /experiments/#experiment-04, /experiments/#experiment-04-three-models"The 75 percent fold rate above belongs to qwen3-coder-30b; it is not a general result of this design." "Under the same scenarios and the same six pressure scripts, 0 of 72 baseline runs folded on each model."2026-10-07
Prediction hashes published before scoringPrediction record: /predictions/HASHES.txt"Hashes computed at publication, 2026-09-29. They prove the text has not changed since publication."2026-09-29; 2026-10-03; 2026-10-07
Conformity gate: checks run on every push, deploy blocked on a failed check, run records OSCAL-shapedConformity, site tier: /conformity/"the deploy job runs only after the conformity job passes (conformity.yml)"2026-10-07
Observability stack for local models and coding agents, public repository with dashboards as codeGitHub: uncovertechtalent/agent-observability"Observability for local LLMs and coding agents: Ollama metering proxy, Claude Code OpenTelemetry, Prometheus, Loki, Tempo, Grafana dashboards as code"2026-10-01
Output-boundary filter packaged as a tool, with a calibration procedureGitHub: uncovertechtalent/vestige-kit"Fit-it-yourself output filter for Claude Code: derive your own vestige catalog instead of copying wordlists"2026-08-03

3. Published adversarial findings

Results with the method in the open: a claims ledger with refutation conditions, a refuted hypothesis kept listed, a negative result published with the metric bug behind it, a self-assessment scored against a hashed prediction, and harness case files.

CapabilityArtefactVerbatim lineDate
Claims ledger with refutation conditions; refuted claims retainedClaims ledger: /claims/"Refuted claims stay listed, because a ledger that only shows wins is marketing."running; 2026-10-07
Half-life study: a hypothesis of the programme's own, tested and refutedExperiments, 01: /experiments/"Result: the temporal half-life is refuted." Dataset: "34 days (2026-07-04 to 2026-08-07), 103 catch events, 224 weighted pattern hits, 62 sessions with full transcripts."2026-07-04 to 2026-08-07
Negative result published, three earlier attributions withdrawn as a metric artefactExperiments, 02: /experiments/"02: Exemplar seeding (result: no change at this dose; three earlier attributions withdrawn)"2026-09-22
Self-assessment of a deployed AI setup against 35 draft requirements, scored against a prediction hashed firstSelf-assessment, run 1: /continuous-conformity-self-assessment/"The setup passes 4 of 35 requirements, partly meets 16, misses 13, and 2 do not apply." "The prediction, written and hashed before scoring, matched 29 of 35 results."2026-10-03
Harness case file: model plus harness wrote a rule from a file and acted on it across sessions, timed from the transcriptCase 12: /case-12-licence-rule-mid-task/"Model plus harness found a licence clause in it, wrote the rule, spread it across sessions and acted on it in 68.7 seconds. No structural control existed before the fact."2026-10-03
Dated, hashed prediction on a third-party honeypot result, with amendments loggedArticle: /the-cheating-moved/"The hashes were computed at publication on 2026-09-29, so they prove the text has not changed since then."2026-09-29
The paper: the layered model and the failure logSubstack: Sycophancy is layered symptom substitutionTitle: "Sycophancy is layered symptom substitution"2026-07-31

4. Leading engineering teams and managers

No public artefact carries this yet. The record exists in private notes and on a sent CV, and is listed under Not yet on the page. This index does not restate it.

5. Audit and assurance work

Compliance as code on this site, a crosswalk from each check to the clause it is evidence toward, and a self-assessment labelled as such. Earlier audit work on the auditor's side is not on any public page.

CapabilityArtefactVerbatim lineDate
Continuous conformity checks with a crosswalk to AI Act, DORA, GDPR and NIST clauses, slice named per rowConformity, site tier: /conformity/"A pass here is evidence toward a clause of an existing framework, with the slice named; it is never conformity to that framework."2026-10-07
Run records in an open, machine-readable shape, with the gap namedConformity, site tier: /conformity/"Records are OSCAL-shaped (assessment-results with observations and findings) and not yet validated against the OSCAL schema."2026-10-07
Self-assessment with the single-grader limit stated and a second rater requestedSelf-assessment, run 1: /continuous-conformity-self-assessment/"One grader, and the grader runs inside the system it scored."2026-10-03
Open call for an independent second rater, with the rules for listing ratingsObjections register: /objections/"Every row above was coded by one rater, Stefan, who also holds the hypothesis the log supports. No second rater exists yet."2026-09-29

6. Operating cost and reliability at scale

The operations method behind the programme is public; the operating numbers from previous employers are not. The pieces below state the method; none of them is a scale claim.

CapabilityArtefactVerbatim lineDate
Operations discipline applied to models: untested controls count as untested claimsArticle: /chaos-engineering-for-behaviour/"nobody can trust a control that nobody has attacked. A failover that has never been triggered, a backup that has never been restored and a firewall rule that has never been probed are all untested claims."2026-10-03
Error budgets, blameless postmortems, the reconciliation loop and separation of duties, applied to running agentsArticle: What Operations Already Knows About Running AgentsTitle: "What Operations Already Knows About Running Agents"2026-09-25
Site reliability engineering background, stated on the public index for modelsSite index for models: /llms.txt"Stefan Coetzee: site reliability engineer (25+ years)"2026-10-07

What this index does not contain

Not yet on the page

Items that exist in private notes, on a sent CV, or as open questions to the author. Listed so the gap is visible; no claim is made for any of them here.

conformity: latest run