arXiv:2605.18915cs.CRcs.AI2026-05ACL

用多图分发指令,突破多模态大模型安全防线

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

论文配图:DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
图 1 · 摘自论文原文
  • 将攻击指令拆分到多张图片中,结合视觉证据与数字链任务
  • 在GPT-4o等模型上实现超90%攻击成功率,远超单图方法
  • 揭示多图像输入下安全对齐的深层缺陷,适合安全研究者参考

多模态大语言模型(MLLMs)易受越狱攻击,可能引发有害响应。许多MLLM支持多图像输入,但针对多图像的安全对齐工作不足,引入了新漏洞。现有越狱方法仅使用单张图像,限制了攻击空间:无法跨图分布指令、传递复杂信息,也无法利用额外视觉推理任务干扰模型。为此,本文提出组合式越狱框架DMN,融合分布式指令、多模态证据与数字链任务,全面增强越狱效果。大量实验表明,DMN在GPT-4o、Gemini-2.5-pro和Claude Sonnet 4上攻击成功率均超过90%,显著优于基线方法。该多图像组合策略暴露了其安全机制的根本弱点。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-image inputs, inadvertently introducing new vulnerabilities due to less efforts on multi-image safety alignment. Previous MLLM jailbreak methods only uses a single image, which restricts the attack space: they cannot distribute harmful requests across multiple images, carry abundant information, or exploit additional visual reasoning tasks to distract MLLMs. To address these limitations, in this paper, we propose a compositional jailbreak framework, \textbf{DMN}, which leverages \textbf{D}istributed instruction, \textbf{M}ultimodal evidence and a \textbf{N}umber chain task to fully enhance the jailbreak performance. Extensive experiments show that DMN is highly effective for MLLM jailbreaking, e.g. achieving attack success rates of over 90\% on GPT-4o, Gemini-2.5-pro and Claude Sonnet 4, surpassing other baselines by a large margin. This compositional, multi-image jailbreak strategy reveals fundamental weaknesses in their safety mechanisms.

越狱攻击多模态安全图像注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。