用视觉干扰破解多模态大模型安全机制,成功率超七成。
Distraction is All You Need for Multimodal Large Language Model Jailbreaking
- 通过拆分提示词和构造对比子图,制造多层级干扰
- 在五类场景中平均攻击成功率达52.40%,集成攻击达74.10%
- 揭示注意力分散是突破模型防护的新路径,适合安全研究者
多模态大语言模型(MLLMs)连接视觉与文本数据,推动诸多应用发展。然而,视觉元素间复杂交互及其与文本对齐可能引入漏洞,被用于绕过安全机制。我们分析图像内容与任务关系,发现子图像复杂度而非内容是关键因素。据此提出干扰假说,并构建对比子图干扰越狱框架(CS-DJ),通过两级策略实现越狱:结构化干扰通过查询分解引发分布偏移,将有害提示拆分为子查询;视觉增强干扰通过构造对比子图破坏模型内部视觉元素交互。该双重策略分散模型注意力,削弱其检测和抑制有害内容能力。在五个代表性场景及四款主流闭源模型(GPT-4o-mini、GPT-4o、GPT-4V、Gemini-1.5-Flash)上的实验表明,CS-DJ平均攻击成功率达52.40%,集成攻击成功率达74.10%。结果揭示基于干扰的越狱方法潜力,为攻击策略提供新视角。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model's attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs' defenses, offering new insights for attack strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。