arXiv:2605.10764cs.CVcs.AI2026-05

通过最大化决策词熵,实现跨模型通用越狱攻击。

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

论文配图:Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
图 1 · 摘自论文原文
  • 在生成过程中放大高熵词的不确定性以触发越狱。
  • 在三个模型上成功率超基准方法,且跨模型迁移能力显著提升。
  • 适合研究模型安全与对抗攻击的学者参考。

近期研究表明,基于梯度的视觉-语言模型(VLMs)通用图像越狱攻击几乎无跨模型可迁移性,引发对可迁移多模态越狱可行性的质疑。我们在严格无目标威胁模型下重新审视这一结论,不强制固定前缀或响应模式。初步实验发现,拒绝行为集中在自回归解码中的高熵词上,而非拒绝词在攻击前已占据顶级候选概率质量。受此启发,我们提出无目标越狱熵最大化方法(UJEM-KL),通过最大化这些决策词的熵来翻转拒绝结果,同时稳定低熵位置以保持输出质量。在三个VLM和两个安全基准上,UJEM-KL达到有竞争力的白盒攻击成功率,并持续提升迁移能力,且对典型防御仍有效。实验表明,有限迁移性主要源于优化目标过度受限。

原文摘要 · Abstract (English)

Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.

越狱攻击视觉语言模型熵最大化对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。