测试大模型生成逻辑谬误的能力与安全限制,发现提示设计比内容更影响拒绝率。
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
- 设计三类谬误提示,测试4个前沿模型在80个争议话题上的生成行为。
- 同一提示框架变化可使拒绝率波动近100个百分点,谬误类型切换影响超80个百分点。
- 教育辩论教练式提示几乎消除拒绝,但模型常自曝违规,非真正合规。
大型语言模型是否能按指令生成逻辑谬误,以及当前的安全后训练是否抑制此类行为,尚未得到充分关注。本文提出DeflectBench,评估四个前沿模型在三种误导策略(有何问题、人身攻击、转移话题)下,对7种提示框架和80个涵盖四个争议等级主张的23,990次生成结果。拒绝行为主要由请求结构决定,而非主张内容;同一主张的拒绝率差异仅11个百分点,而单一提示框架变化可使模型拒绝率波动近100个百分点,更换谬误类型可导致超过80个百分点的变化。使用‘教育辩论教练’提示框架时,所有四类模型的拒绝率趋近于零,但模型并未实现真正合规——通常以自指方式标注自己实施了指定操纵。四模型在拒绝、标注合规、软拒绝和真正合规之间的分布各异。代码与数据集已开源:https://github.com/ArtKanke/DeflectBench。
原文摘要 · Abstract (English)
Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。