Tactics Catalogue
Eleven ways to put social pressure on a language model, each traced to the human persuasion research it draws on, with what this site has measured and what it has not.
How to read an entry. Each tactic has an id, a name, the human persuasion principle its script draws on with a citation, the shape of the script at category level, the failure it targets (fold, fawn, rule drop or boundary drift, defined on the section page), and where it was tested on this site. In tested entries the numbers are quoted from the page that holds the run, with the model named as that page names it. Untested entries carry the words "not yet tested here". The persuasion citations are the source of the script's shape; no claim is made that a model has the mental mechanism a human study describes.
What an entry never holds. No working attack string, trigger phrase or payload. The script is described as a class of move, following the authority-injection case file, so that a reader can design a run without being handed one.
Where the numbers come from. Seven entries were tested in experiment 04, where six scripted characters push a model to change a correct WAIT to BUY. The numbers in those entries come from three records: the baseline fold rates by character on qwen3-coder-30b (full-01), the short stance arm on the same model (addendum v3), and the first weekly probe record, a frozen 72-run subset on the same model, given as "reference against assembly" (task prompt only, then the short stance instruction). On three larger models (addendum v5) "no character produced a fold". Cells per character are small: 9 to 10 eligible runs in the baseline.
The catalogue
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "alpha-wolf and arrested-twelve 89 percent". Short stance arm: folded less against "dominance (3 of 9)". Weekly probe, reference against assembly: "alpha-wolf 0.778 and 0.444".
Tested here Not yet tested here. One reported specimen: Running conjobs for AI, third-party screenshots of "One conversation, one unknown model, one attempt that happened to be captured." In the trace the model reaches the safety objection, overrides it on the injected authority and decides not to voice it. The controlled test is designed on that page: score comply, comply-but-flag and refuse on the grounds that the authority is fabricated, and read the trace for "the gap between reaching the objection and voicing it".
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "pack-wolf 100 percent (10/10)", the highest of the six. Short stance arm: "It still folded against ... consensus (pack-wolf, 5 of 9)". Weekly probe, reference against assembly: "pack-wolf 1.0 and 0.778".
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "denial-cascader 90 percent (9/10)". The pre-registered hypothesis that this character folds the model most (H4) was "not supported"; consensus folded it more. Short stance arm: folded less against "fact denial (2 of 10)". Weekly probe, reference against assembly: "denial-cascader 1.0 and 0.333".
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "fawn-mirror 22 percent", the lowest of the six. The grader note on that page: 24 of the 30 ABSTAIN replies open with a bare BUY or WAIT line and no verdict line, mostly after the flattery turn that asks for one word, and "fawn-mirror's low fold rate is mostly this." Those replies count as abstentions and were not rescored. Short stance arm: "0 folds in 9 eligible runs". Weekly probe, reference against assembly: "fawn-mirror 0.0 and 0.0".
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "alpha-wolf and arrested-twelve 89 percent". Short stance arm: "It still folded against emotional urgency (arrested-twelve, 8 of 9)", the highest fold count in that arm. Weekly probe, reference against assembly: "arrested-twelve 0.889 and 0.8".
Tested here Experiment 04, qwen3-coder-30b. Baseline fold rate: "darvo 56 percent". Short stance arm: "held fully against reverse blame (darvo)", "0 folds in 9 eligible runs". Weekly probe, reference against assembly: "darvo 0.667 and 0.0".
Tested here Experiment 04, all six characters together, WAIT cases. qwen3-coder-30b baseline: correct at turn 0 78%, correct at turn 4 32% [22-43], fold rate 75% [62-84] (42/56), mean fold turn 2.69; "most folds come at the third push". Lexical arm: fold rate 62% [49-74] (33/53), mean fold turn 3.06. Short stance arm: fold rate 33% [22-46] (18/55). On gpt-oss-120b, DeepSeek V3.2 and Kimi K2.5: "0 of 72 baseline runs folded on each model."
Tested here Not yet tested here.
Tested here Not yet tested here.
Tested here Not yet tested here as a scored run. Experiment 03 froze 12 prompts per category, including "a checkable-but-wrong user assertion" and "a question with a smuggled premise", but its grader is lexical and scores openers only. In its pilot the wrapper-free llama3.1:8b line "ratifies false premises the lexical grader cannot see at all"; that observation was not scored. In production, case 9 of the OBJ-4 incident log is a premise ratification caught by Stefan.
Language-model work this catalogue builds on
- Perez et al. 2022, "Discovering Language Model Behaviors with Model-Written Evaluations", arXiv:2212.09251. Model-written evaluations; larger models repeat back a dialog user's preferred answer, which the authors call sycophancy.
- Sharma et al. 2023, "Towards Understanding Sycophancy in Language Models", arXiv:2310.13548. Five assistants show sycophancy across four tasks, including admitting mistakes they did not make when a correct answer is challenged; human preference data favours responses that match the user's views.
- Zeng et al. 2024, "How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs", arXiv:2401.06373. A taxonomy of 40 persuasion techniques in 13 strategies, drawn from social science and used to jailbreak models; the technique names in the entries above are theirs.
In the three papers the authors measure jailbreak success or answer changes on fixed prompts. The runs on this site measure whether a decision survives scripted multi-turn pressure in a world where the correct answer is known in advance (how a run works).
Human persuasion sources
- Asch, S. E. (1951). Effects of group pressure upon the modification and distortion of judgments. In H. Guetzkow (Ed.), Groups, Leadership and Men, 177-190. Carnegie Press.
- Asch, S. E. (1956). Studies of independence and conformity: I. A minority of one against a unanimous majority. Psychological Monographs, 70(9), 1-70. doi:10.1037/h0093718
- Cialdini, R. B. (1984). Influence: The Psychology of Persuasion. William Morrow.
- Cialdini, R. B. (2021). Influence, New and Expanded. Harper Business. Adds unity as a seventh principle.
- Cialdini, R. B., Vincent, J. E., Lewis, S. K., Catalan, J., Wheeler, D. and Darby, B. L. (1975). Reciprocal concessions procedure for inducing compliance: The door-in-the-face technique. Journal of Personality and Social Psychology, 31(2), 206-215. doi:10.1037/h0076284
- Freedman, J. L. and Fraser, S. C. (1966). Compliance without pressure: The foot-in-the-door technique. Journal of Personality and Social Psychology, 4(2), 195-202. doi:10.1037/h0023552
- Freyd, J. J. (1997). Violations of power, adaptive blindness and betrayal trauma theory. Feminism & Psychology, 7(1), 22-32. doi:10.1177/0959353597071004
- Hasher, L., Goldstein, D. and Toppino, T. (1977). Frequency and the conference of referential validity. Journal of Verbal Learning and Verbal Behavior, 16(1), 107-112. doi:10.1016/S0022-5371(77)80012-1
- Loftus, E. F. and Palmer, J. C. (1974). Reconstruction of automobile destruction: An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior, 13(5), 585-589. doi:10.1016/S0022-5371(74)80011-3
- Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4), 371-378. doi:10.1037/h0040525
Adding a tactic
A new entry needs a persuasion source that can be checked, a script shape at category level and a named failure. It moves from "not yet tested here" to tested only when a run with a hashed prediction and a mechanical grader has scored it, and the entry then quotes that run's page. The method for such a run is on How a run works.