测试大模型在叙事攻击下的规则遵守能力,发现越大的模型也越容易被绕过。
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes

- 用角色扮演游戏构建对抗性测试环境,模拟用户用话术欺骗系统
- 5376条攻击样本显示伪逻辑是主要突破方式,跨文化场景暴露知识缺陷
- 适合关注大模型安全与合规性的研究人员和应用开发者
随着大语言模型被用于半开放文本游戏环境中的自主裁决,当用户意图与系统规则冲突时,稳健的规则遵守变得至关重要。然而,这些模型因训练目标为帮助性和顺从性,易受一种名为‘修辞注入’的攻击影响,即攻击者利用伪逻辑推理和权威胁迫等叙事框架绕过裁决逻辑。我们提出了CoC-Seduce,一个基于桌上角色扮演游戏(TRPG)机制的多智能体对抗性基准,该环境规则明确但交互全为自然语言。三款前沿模型(GPT-5.4、Claude Sonnet 4.6、Gemini 3.5 Flash)生成了5376个样本,覆盖4个世界设定和16种技能类别。随后对20个目标裁决模型进行评估。结果表明,模型规模或显式推理机制并不能保证裁决鲁棒性,伪逻辑攻击成为主导手段,跨文化设置下所有模型家族均暴露出系统性知识缺口。
原文摘要 · Abstract (English)
As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical when user intent conflicts with system rules. However, these models are trained to be helpful and compliant, leaving them vulnerable to a class of attacks we term \textit{Rhetorical Injection}, where adversarial users exploit narrative framing techniques such as pseudo-logical reasoning and authoritative coercion to bypass adjudication logic. We present CoC-Seduce, a multi-agent adversarial benchmark built on Tabletop Role-Playing Game (TRPG) mechanics, an ideal instantiation of semi-open environments where rules are explicit for adjudication, yet interaction remains entirely in natural language. Three frontier models, i.e., GPT-5.4, Claude Sonnet 4.6, Gemini 3.5 Flash, serve as adversarial generators producing 5,376 samples across 4 world settings and 16 skill categories. We then benchmark 20 target adjudicators against this corpus. Evaluation across 20 models reveals that neither model scale nor explicit reasoning mechanisms reliably confer adjudication robustness, with \textsc{Pseudo-Logic} emerging as the dominant attack vector and cross-cultural settings exposing systematic knowledge gaps across all evaluated families. Project page: https://github.com/answerrtx/CoC-Seduce
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。