arXiv:2606.03647cs.CRcs.AI2026-06

提出新型黑箱攻击方法,可高效突破大模型安全防护。

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

论文配图:Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs
图 1 · 摘自论文原文
  • 用迭代偏好优化训练掩码扩散语言模型作为攻击者
  • 在无防御适配下对多类防御系统攻击成功率显著提升
  • 适合评估大模型安全性,尤其适用于未公开模型的测试

准确评估对抗鲁棒性长期面临挑战。错误的攻击设计会虚高鲁棒性估计,导致部署风险评估与防御对比不可靠。图像分类器已有AutoAttack等标准化攻击提供可靠基准,但大模型越狱评估尚无类似工具,且设计难度更高。一个可靠的攻击需具备黑箱、可迁移、通用、高效等特性,而现有方法无法同时满足。本文提出间接有害性优化(IHO),一种通过与有害性判别器进行迭代偏好优化训练的掩码扩散语言模型攻击者,仅需目标模型的黑箱访问。该方法无需修改即可作为强自适应攻击,或作为可迁移的通用策略,应用于未见行为和未见目标模型,无需微调。即使面对分层防御(如电路断路器训练模型+辅助检测器),IHO仍显著优于现有最先进方法,且无需任何针对防御的适配。结果表明,IHO是迈向标准化越狱评估的重要一步。代码与模型已在GitHub和Hugging Face开源。

原文摘要 · Abstract (English)

Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably more difficult. A reliable attack must, among other things, be black-box compatible, applicable to arbitrary defense pipelines, and efficient, which no existing method jointly satisfies. We introduce Indirect Harm Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requiring only black-box access to the target. The same method can be used without modification as a strong adaptive attack on individual behaviors, or as an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning. Even against layered defenses, such as a Circuit Breaker-trained model combined with an auxiliary detector, IHO improves attack success considerably over state-of-the-art approaches, without any defense-specific adaptation. Our results position IHO as a practical step toward the kind of standardized jailbreak evaluation that has improved reliability in the past. Code and models are available on GitHub and Hugging Face.

大模型安全越狱攻击黑箱攻击扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。