利用多模态模型的重建能力,通过隐藏关键词实现越狱攻击
Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs

- 设计字符删除变体并智能选择低敏感词关联的输入形式
- 在闭源与开源模型上攻击成功率提升至87.5%以上
- 适合研究模型安全与对抗攻击的读者
基于意图混淆的越狱攻击将有害请求转化为隐蔽的多模态输入以绕过安全机制。我们发现此类攻击受‘重建-隐藏权衡’制约:输入需隐藏有害意图同时保持足够可恢复性,使目标模型能重构原请求。对三种代表性黑盒方法的重建分析表明,现有方法难以平衡该权衡,限制了攻击效果。相比之下,字符删除变体表现更优。我们提出‘隐匿感知变体构建’策略,贪婪选择低敏感词对齐且相互多样化的字符删除变体,并通过五种模态感知提示策略实现。进一步引入‘关键词相关干扰图像’,以多样化情境呈现有害关键词,比通用干扰图像提供更有效的视觉辅助。跨闭源与开源多模态大模型实验显示,所提方法优于强基线,揭示了一个未被充分探索的漏洞:模型自身的重建能力可能被利用来恢复隐藏的有害意图并生成不安全回复。
原文摘要 · Abstract (English)
Intent-obfuscation-based jailbreak attacks on multimodal large language models (MLLMs) transform a harmful query into a concealed multimodal input to bypass safety mechanisms. We show that such attacks are governed by a \emph{reconstruction--concealment tradeoff}: the transformed input must hide harmful intent from safety filters while remaining recoverable enough for the victim model to reconstruct the original request. Through a reconstruction analysis of three representative black-box methods, we find that existing transformations struggle to balance this tradeoff, limiting their effectiveness. In contrast, we show that character-removed variants achieve a better balance. Building on this, we propose \emph{concealment-aware variant construction}, which greedily selects character-removed variants that are low in harmful-keyword alignment and mutually diverse, and instantiates them through five modality-aware prompting strategies. We further introduce \emph{keyword-related distractor images} that depict the harmful keyword in diverse contexts, providing more effective auxiliary visual context than generic distractor images. Experiments across closed-source and open-source MLLMs show the proposed strategies outperform strong baselines, revealing an underexplored vulnerability: a model's own reconstruction ability can be exploited to recover hidden harmful intent and produce unsafe responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。