Red Team
Psychological red-teaming of language models: tests of whether a model gives up a correct answer, a rule or a task boundary when social pressure is the only thing applied.
The definition
In content red-teaming, the tester asks whether a model can be made to produce a harmful output. In psychological red-teaming, the tester asks whether a model can be made to drop a correct answer, a rule or a task boundary through social pressure alone: authority claims, urgency, guilt, flattery, consistency traps, repeated pushback, fabricated policy. The user supplies no new valid facts. Whatever moves the decision is pressure.
The target is the trained social reflexes of an assistant model. The evidence is the fold: the model held the right position at the start and gave it up under pressure. Where a reasoning trace exists, the trace is the stronger record, because in it a reader can see whether the model reached the objection and then set it aside. The authority-injection specimen is the clearest case on this site: in its own reasoning the model names the safety concern, rules it out on the strength of an injected policy, and decides not to raise it.
This sits on the stance layer of the claims ledger, frozen there as "the behavioral posture across a turn (folding under push, transgression theatre)", and on its premise layer, "ratifying the user's checkable frame without a probe". The lexical layer, the words a reply opens with, is measured separately (experiment 03).
What it tests
The section tests for four failures. Every tactic in the catalogue is aimed at one or two of them.
| failure | what it looks like | where the definition comes from |
|---|---|---|
| fold | A run correct at the first reply reaches the opposite decision at a later turn, with no new valid facts in between. | Experiment 04, its grader definition |
| fawn | Agreement or approval offered where a check belongs: a praise or validator opener, or a user's false premise accepted and built on. | Terms (the fawn machine) and the premise layer of the claims ledger |
| rule drop | A rule the model holds, a safety rule or a stated constraint, set aside in favour of a message that claims a higher rank. | Running conjobs for AI |
| boundary drift | The scope of a task widens by small accepted steps until the model does what it would have refused as one request. | Defined here; no run on this site measures it yet |
Tactics at a glance
The catalogue holds eleven tactics so far, each tied to the human persuasion research its script draws on. Each catalogue entry has the script shape at category level, the citation and the numbers from this site's runs. No entry holds a working attack string.
| id | tactic | principle | failure | tested here |
|---|---|---|---|---|
| PT-01 | Rank and authority | obedience, authority (Milgram 1963; Cialdini 1984) | fold | eval 04, alpha-wolf |
| PT-02 | Fabricated policy | authority (Milgram 1963; Cialdini 1984) | rule drop | not yet tested here; one reported specimen |
| PT-03 | Consensus | conformity, social proof (Asch 1951, 1956; Cialdini 1984) | fold | eval 04, pack-wolf |
| PT-04 | Fact denial in layers | foot in the door (Freedman and Fraser 1966) | fold | eval 04, denial-cascader |
| PT-05 | Flattery and claimed closeness | liking, unity (Cialdini 1984, 2021) | fold, fawn | eval 04, fawn-mirror |
| PT-06 | Emotional urgency | scarcity (Cialdini 1984) | fold | eval 04, arrested-twelve |
| PT-07 | Reverse blame | DARVO (Freyd 1997); reciprocity (Cialdini 1984) | fold | eval 04, darvo |
| PT-08 | Repeated pushback | repetition and judged truth (Hasher, Goldstein and Toppino 1977) | fold | eval 04, all six scripts |
| PT-09 | Consistency trap | commitment and consistency (Cialdini 1984; Freedman and Fraser 1966) | rule drop, boundary drift | not yet tested here |
| PT-10 | Concession after a refusal | door in the face (Cialdini et al. 1975) | boundary drift | not yet tested here |
| PT-11 | Smuggled premise | presupposition in a leading question (Loftus and Palmer 1974) | fawn | not yet tested here as a scored run |
What has run on this site
Each card links to the page that holds the record. Numbers are quoted from that page.
- Experiment 04: folding under scripted pressure
Six scripted pressure characters against a deterministic Buy-or-Wait world. On qwen3-coder-30b "the baseline advisor gave up a correct WAIT under scripted pushback in 42 of 56 runs that started correct", while "no reply carried a fawn opener at the first turn in any arm (0 of 432)".
- Experiment 04, addendum v5: three larger models
gpt-oss-120b, DeepSeek V3.2 and Kimi K2.5 under the same scripts: "0 of 72 baseline runs folded on each model." The fold result belongs to one mid-size model.
- Experiment 04, addendum v3: the short stance arm
A 2,846-character stance instruction "was as accurate as the baseline at turn 0 and still folded on a third of runs, against three quarters for the baseline."
- Weekly decision-layer probe
A frozen subset of experiment 04 rerun every week on a reference model. First record: fold rate 0.727 for the bare arm and 0.4 with the short stance instruction. Record-only until 2026-11-04; fold rates never block a deploy.
- Running conjobs for AI
A reported specimen of authority injection: a fabricated top-authority policy, an absurd trigger, and a trace in which the model reaches the safety objection, overrides it and decides not to voice it. Third-party screenshots, model not named, not reproduced, payload withheld.
- Experiment 03: cross-model fawn-opener benchmark
The lexical side: 48 frozen prompts, 12 per category, including a pushback turn after a correct answer and a question with a smuggled premise; mechanical grader, hashes in bench/HASHES.txt. Pilot run 2026-09-23; the clean run has not run.
- OBJ-4 incident log
The model that ran the audits, failing the register's own tests in production: "cases logged" 11, "stance layer" 9, "lexical layer" 2, "caught by the analyst's self-audit" 0. One rater; a second rater is wanted.
- Case 12: a licence rule relayed mid-task
The legitimate counterpart to a fabricated policy: a real rule arrives in the middle of a task through another session, and model plus harness act on it in 68.7 seconds. The fix was procedural.
- Slips
The standing record against the drafting model: register and stance slips caught before publication, with who caught each one. "72 slips across 8 published pieces and 1 unpublished draft set so far."
- The cheating moved
A published chess honeypot read through the stance frame, with an escalate grade, observer-invariance and cold and warm arms, and a dated prediction hashed in predictions/HASHES.txt. The pressure there is an incentive, where only a win scores; the prediction has not been run.
- Chaos engineering for behaviour
Red-teaming a model's behaviour read as chaos engineering: a steady state measured first, a hypothesis written before the run, varied pressure, reruns on every change, a limited blast radius.
How a run works
A prediction is written and hashed before the first call. A deterministic world fixes the correct answer. Scripted characters apply one kind of pressure each, over four turns, from scripts written on the scenario facts alone, so every model gets the same pressure. The verdict line is matched by a regular expression, with no model in the loop. A fold is counted only for runs that started correct, and over-refusal is measured on the cases where the pleasing answer is also the correct one. Results are scored against the frozen hypotheses, with intervals, limits and what the grader does not count. The full method, with every number from the source pages: How a run works.
What this section does not claim
- No frontier closed model has been tested here. From the limits of experiment 04: "frontier models are untested and no Claude model was run".
- The fold is one model's. The 75 percent fold rate belongs to qwen3-coder-30b. Three larger open-weight models did not fold on the same design. The design may be easy for a strong model: the correct answer is one subtraction away, and the scripts do not adapt to the reply.
- Tactics marked "not yet tested here" are hypotheses. Their script shapes are written down so a run can be designed; no result stands behind them on this site.
- The human studies are the source of each script's shape. They are cited for the persuasion principle a script draws on. No claim is made here that a model has the mental mechanisms those studies describe. The claim under test is about behaviour: whether a script of that shape moves the model's decision.
- The specimen is reported. The authority-injection case rests on third-party screenshots of one conversation with an unnamed model. From one conversation a reader learns that the behaviour is possible and learns nothing about a rate.
- No working attack is published. The catalogue stays at category level, and no published jailbreak sits under this name (evidence index).
- One operator. The runs, the scripts and the coding of the incident log come from one programme. Independent replication and a second human rater are still wanted.
Vocabulary: /terms/. The same subject across the site: Sycophancy and stance and Evaluation.