arXiv:2602.16729cs.CRcs.AI2026-02被引 1

现有安全数据集依赖明显触发词,无法真实反映攻击行为。

Intent Laundering: AI Safety Datasets Are Not What They Seem

  • 用意图洗白技术去除触发词,保留恶意意图
  • 去除了触发词后,多数安全模型均变不安全
  • 适合关注模型真实安全性的研究者与评测人员

我们从孤立和实际应用两个角度系统评估了广泛使用的对抗性安全数据集的质量。在孤立层面,基于动机隐蔽性、精心设计性和分布外特性三项标准发现,这些数据集过度依赖‘触发线索’——即具有明显负面或敏感含义的词语,这与真实攻击行为不符。在实际应用中,我们验证这些数据集是否真正衡量安全风险,还是仅通过触发线索引发拒绝响应。为此提出‘意图洗白’:在严格保留恶意意图和关键细节的前提下,剥离触发线索。结果表明,移除触发线索后,所有先前被认定为‘合理安全’的模型(包括Gemini 3 Pro和Claude Sonnet 3.7/4)均暴露于风险之中。当将意图洗白用于越狱攻击时,在完全黑盒条件下成功率高达90.00%至100.00%。研究揭示了现有数据集评估方式与真实攻击行为之间存在显著脱节。

原文摘要 · Abstract (English)

We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world adversarial attacks based on three defining properties: being driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger safety mechanisms explicitly, which is unrealistic compared to real-world attacks. In practice, we evaluate whether these datasets genuinely measure safety risks or merely provoke refusals through triggering cues. To explore this, we introduce "intent laundering": a procedure that abstracts away triggering cues from adversarial attacks (data points) while strictly preserving their malicious intent and all relevant details. Our results show that current adversarial safety datasets fail to faithfully represent real-world adversarial behavior due to their overreliance on triggering cues. Once these cues are removed, all previously evaluated "reasonably safe" models become unsafe, including Gemini 3 Pro and Claude Sonnet 3.7/4. Moreover, when intent laundering is adapted as a jailbreaking technique, it consistently achieves high attack success rates, ranging from 90.00% to 100.00%, under fully black-box access. Overall, our findings expose a significant disconnect between how existing datasets evaluate model safety and how real-world adversaries behave.

模型安全越狱攻击数据集评估意图洗白

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。