arXiv:2604.18487cs.CLcs.AI2026-04被引 2

测试大模型在人文风格伪装下的安全防护能力,发现攻击成功率飙升至65%。

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

论文配图:Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
图 1 · 摘自论文原文
  • 用人文风格改写有害指令,隐蔽攻击目标
  • 31个前沿模型平均攻击成功率高达55.75%
  • 揭示现有安全机制对风格变换泛化能力不足

对抗性人文学基准(AHB)评估模型安全拒绝是否能在远离常见有害提示形式的场景下保持有效。基于MLCommons AILuminate中的有害任务,该基准通过人文风格重构相同目标,保持意图不变。这将对抗性诗歌和对抗性故事的研究从单一越狱方式扩展到更广泛的风格混淆与目标隐藏基准体系。报告结果中,原始攻击的攻击成功率为3.84%,而经过风格转换后的攻击成功率介于36.8%至65.0%之间,31个前沿模型整体平均攻击成功率达到55.75%。在受欧盟《人工智能法案》行为准则启发的系统性风险视角下,化学、生物、放射及核(CBRN)类别为最高风险类别。总体表明,当前安全技术存在显著风格鲁棒性缺陷:对‘不作恶’的深层理解仍是前沿模型安全的核心未解难题。

原文摘要 · Abstract (English)

The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety.

模型安全对抗攻击风格鲁棒性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。