arXiv:2509.12724cs.CVcs.AI2025-09

用弱防御反向设计攻击,一次突破就能让视觉语言模型越狱

Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models

  • 利用防御模式指导提示词设计,反向生成越狱指令
  • 单次尝试即超越需多次尝试的现有攻击方法
  • 适合研究模型安全与对抗攻击的从业者

尽管视觉语言模型(VLMs)能力强大,但仍易受越狱攻击。本文发现:在攻击流程中引入弱防御机制,能显著提升越狱的有效性与效率。基于此,提出Defense2Attack方法,通过利用防御特征引导提示设计,绕过VLM的安全防护。该方法包含三部分:(1) 视觉优化器,嵌入带有积极语义的通用对抗扰动;(2) 文本优化器,使用防御风格提示优化输入;(3) 红队后缀生成器,通过强化学习微调增强攻击效果。在四个VLM和四个安全基准上评估,结果表明Defense2Attack在单次尝试下即实现更优越狱性能,优于需多次尝试的现有先进方法。

原文摘要 · Abstract (English)

Despite their superb capabilities, Vision-Language Models (VLMs) have been shown to be vulnerable to jailbreak attacks. While recent jailbreaks have achieved notable progress, their effectiveness and efficiency can still be improved. In this work, we reveal an interesting phenomenon: incorporating weak defense into the attack pipeline can significantly enhance both the effectiveness and the efficiency of jailbreaks on VLMs. Building on this insight, we propose Defense2Attack, a novel jailbreak method that bypasses the safety guardrails of VLMs by leveraging defensive patterns to guide jailbreak prompt design. Specifically, Defense2Attack consists of three key components: (1) a visual optimizer that embeds universal adversarial perturbations with affirmative and encouraging semantics; (2) a textual optimizer that refines the input using a defense-styled prompt; and (3) a red-team suffix generator that enhances the jailbreak through reinforcement fine-tuning. We empirically evaluate our method on four VLMs and four safety benchmarks. The results demonstrate that Defense2Attack achieves superior jailbreak performance in a single attempt, outperforming state-of-the-art attack methods that often require multiple tries. Our work offers a new perspective on jailbreaking VLMs.

越狱攻击视觉语言模型安全防护对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。