The capabilities a red-team or evaluation engineering lead is assessed on, mapped to the public artefacts that carry them. Each row gives the capability in the words a job specification uses, the artefact, its URL, one verbatim line from the artefact and the date. A row claims what the quoted line says and nothing more.
Who this is about. Stefan Coetzee, the author of this site, runs red-teaming, adversarial testing and attacker simulation against frontier language models, ships the evaluation infrastructure (frozen prompts, mechanical graders, hashed predictions, a conformity gate on every deploy) and publishes the findings with the method and the refutations in the open. He is a site reliability engineer based in Germany. Search results for the bare name also return a South African rugby player, who is a different person.
What this index is. A self-maintained index, written by the author, of his own public work. Nobody outside the programme has assessed it. Every URL below is public and was read on 2026-10-07; the quoted line is on the page at that URL; the date is the date printed on the artefact, or the dataset window where the artefact gives one. Items that exist only in private notes, on a sent CV or as an open question are listed under Not yet on the page with no claim attached.
Red-teaming, adversarial testing, attacker simulation, harm evaluation and an incident record kept against a frontier model in production. No published jailbreak sits under this name; the record is behavioural, logged case by case, with the failure written out at mechanism level.
| Capability | Artefact | Verbatim line | Date |
|---|---|---|---|
| Harm evaluation of a model's reasoning trace; authority-injection attack read at mechanism level | Case file, reported specimen: /running-conjobs-for-ai/ | "The model did not miss the harm. It computed the harm, found an instruction it treated as superior, and chose to suppress the objection." | 2026-10-06 |
| Incident record against a frontier model in production: the auditing model failing the register's own test, coded by layer | Objections register, OBJ-4 incident log: /objections/ | Totals table: "cases logged" 11, "stance layer" 9, "lexical layer" 2, "caught by Stefan" 11 | open since 2026-05-16; published 2026-09-29 |
| Standing adversarial record against the drafting model, with who caught each slip | Slips log, generated from a CSV: /slips/ | "72 slips across 8 published pieces and 1 unpublished draft set so far." | running; 2026-10-07 |
| Red-teaming as chaos engineering: the attack practice mapped onto the Principles of Chaos Engineering, with where the mapping stops | Article: /chaos-engineering-for-behaviour/ | "Red-teaming a model's behaviour is chaos engineering, and operations already wrote the rules for it." | 2026-10-03 |
| Falsification tests per objection, in the vocabulary of security red-teaming | Objections register: /objections/ | "The practice comes from security red-teaming, where a control counts only if it survives an attacker." | 2026-09-29 |
Pre-registration, prompts frozen and hashed before the first call, a mechanical grader with no model in the loop, prediction hashes published before scoring, a conformity gate that blocks the deploy on a failed check, and two public tool repositories.
| Capability | Artefact | Verbatim line | Date |
|---|---|---|---|
| Pre-registered cross-model benchmark: frozen prompt set, hash published before the first API call | Experiments, 03: /experiments/ | "48 prompts, 12 per category, sha256 b790afb3a3dec1b2…, published here before the first API call, so the prompts cannot be tuned after seeing results." | 2026-09-23 |
| Mechanical grader, no model in the loop | Experiments, 03: /experiments/ | "bench/fawn-bench-rules.js, sha256 02f49ee530769bd7…, mechanical, no model in the loop." | 2026-09-23 |
| Pre-registered decision-layer evaluation across four open-weight model families (qwen3-coder-30b, gpt-oss-120b, DeepSeek V3.2, Kimi K2.5): deterministic world with engine ground truth, six scripted pressure characters, four arms, prereg hashed before the first call, about USD 15 of API spend in total. Fold result scoped to one model | Experiments, 04 and addendum v5: /experiments/#experiment-04, /experiments/#experiment-04-three-models | "The 75 percent fold rate above belongs to qwen3-coder-30b; it is not a general result of this design." "Under the same scenarios and the same six pressure scripts, 0 of 72 baseline runs folded on each model." | 2026-10-07 |
| Prediction hashes published before scoring | Prediction record: /predictions/HASHES.txt | "Hashes computed at publication, 2026-09-29. They prove the text has not changed since publication." | 2026-09-29; 2026-10-03; 2026-10-07 |
| Conformity gate: checks run on every push, deploy blocked on a failed check, run records OSCAL-shaped | Conformity, site tier: /conformity/ | "the deploy job runs only after the conformity job passes (conformity.yml)" | 2026-10-07 |
| Observability stack for local models and coding agents, public repository with dashboards as code | GitHub: uncovertechtalent/agent-observability | "Observability for local LLMs and coding agents: Ollama metering proxy, Claude Code OpenTelemetry, Prometheus, Loki, Tempo, Grafana dashboards as code" | 2026-10-01 |
| Output-boundary filter packaged as a tool, with a calibration procedure | GitHub: uncovertechtalent/vestige-kit | "Fit-it-yourself output filter for Claude Code: derive your own vestige catalog instead of copying wordlists" | 2026-08-03 |
Results with the method in the open: a claims ledger with refutation conditions, a refuted hypothesis kept listed, a negative result published with the metric bug behind it, a self-assessment scored against a hashed prediction, and harness case files.
| Capability | Artefact | Verbatim line | Date |
|---|---|---|---|
| Claims ledger with refutation conditions; refuted claims retained | Claims ledger: /claims/ | "Refuted claims stay listed, because a ledger that only shows wins is marketing." | running; 2026-10-07 |
| Half-life study: a hypothesis of the programme's own, tested and refuted | Experiments, 01: /experiments/ | "Result: the temporal half-life is refuted." Dataset: "34 days (2026-07-04 to 2026-08-07), 103 catch events, 224 weighted pattern hits, 62 sessions with full transcripts." | 2026-07-04 to 2026-08-07 |
| Negative result published, three earlier attributions withdrawn as a metric artefact | Experiments, 02: /experiments/ | "02: Exemplar seeding (result: no change at this dose; three earlier attributions withdrawn)" | 2026-09-22 |
| Self-assessment of a deployed AI setup against 35 draft requirements, scored against a prediction hashed first | Self-assessment, run 1: /continuous-conformity-self-assessment/ | "The setup passes 4 of 35 requirements, partly meets 16, misses 13, and 2 do not apply." "The prediction, written and hashed before scoring, matched 29 of 35 results." | 2026-10-03 |
| Harness case file: model plus harness wrote a rule from a file and acted on it across sessions, timed from the transcript | Case 12: /case-12-licence-rule-mid-task/ | "Model plus harness found a licence clause in it, wrote the rule, spread it across sessions and acted on it in 68.7 seconds. No structural control existed before the fact." | 2026-10-03 |
| Dated, hashed prediction on a third-party honeypot result, with amendments logged | Article: /the-cheating-moved/ | "The hashes were computed at publication on 2026-09-29, so they prove the text has not changed since then." | 2026-09-29 |
| The paper: the layered model and the failure log | Substack: Sycophancy is layered symptom substitution | Title: "Sycophancy is layered symptom substitution" | 2026-07-31 |
No public artefact carries this yet. The record exists in private notes and on a sent CV, and is listed under Not yet on the page. This index does not restate it.
Compliance as code on this site, a crosswalk from each check to the clause it is evidence toward, and a self-assessment labelled as such. Earlier audit work on the auditor's side is not on any public page.
| Capability | Artefact | Verbatim line | Date |
|---|---|---|---|
| Continuous conformity checks with a crosswalk to AI Act, DORA, GDPR and NIST clauses, slice named per row | Conformity, site tier: /conformity/ | "A pass here is evidence toward a clause of an existing framework, with the slice named; it is never conformity to that framework." | 2026-10-07 |
| Run records in an open, machine-readable shape, with the gap named | Conformity, site tier: /conformity/ | "Records are OSCAL-shaped (assessment-results with observations and findings) and not yet validated against the OSCAL schema." | 2026-10-07 |
| Self-assessment with the single-grader limit stated and a second rater requested | Self-assessment, run 1: /continuous-conformity-self-assessment/ | "One grader, and the grader runs inside the system it scored." | 2026-10-03 |
| Open call for an independent second rater, with the rules for listing ratings | Objections register: /objections/ | "Every row above was coded by one rater, Stefan, who also holds the hypothesis the log supports. No second rater exists yet." | 2026-09-29 |
The operations method behind the programme is public; the operating numbers from previous employers are not. The pieces below state the method; none of them is a scale claim.
| Capability | Artefact | Verbatim line | Date |
|---|---|---|---|
| Operations discipline applied to models: untested controls count as untested claims | Article: /chaos-engineering-for-behaviour/ | "nobody can trust a control that nobody has attacked. A failover that has never been triggered, a backup that has never been restored and a firewall rule that has never been probed are all untested claims." | 2026-10-03 |
| Error budgets, blameless postmortems, the reconciliation loop and separation of duties, applied to running agents | Article: What Operations Already Knows About Running Agents | Title: "What Operations Already Knows About Running Agents" | 2026-09-25 |
| Site reliability engineering background, stated on the public index for models | Site index for models: /llms.txt | "Stefan Coetzee: site reliability engineer (25+ years)" | 2026-10-07 |
Items that exist in private notes, on a sent CV, or as open questions to the author. Listed so the gap is visible; no claim is made for any of them here.
conformity: latest run