通过复杂图文叠加突破视觉语言模型安全防线,成功率超80%。
Overloading Large Vision-Language Models for Jailbreaking

- 用递归布局的多维图文组合制造信息过载,干扰跨模态对齐。
- 在开源模型上平均攻击成功率88.6%,商业模型达84.0%,提升48.7%。
- 方法可迁移至不同模型,揭示复杂输入削弱拒绝响应的机制。
大型视觉语言模型(LVLMs)具备强大的跨模态能力,广泛应用于个人助手、文档分析和具身智能体等场景。然而,其双模态输入接口使其易受越狱攻击。现有攻击多依赖简短文本和分布外图像,但近期大语言模型架构与多模态机制的进步已削弱此类攻击的泛化能力。为此,本文提出一种新型信息过载方法,结合大量文本与多维度图像攻击,并采用递归式图文布局,呈指数级提升多模态信息复杂度。该设计显著增加模型跨模态处理负担,破坏其安全对齐机制。在开源与商用LVLM上的实验表明,本方法达到新最佳性能:开源模型平均攻击成功率达88.6%,商用模型达84.0%,较最优基线提升48.7%。此外,基于开源代理模型优化的提示可有效迁移至不同模型家族。进一步分析揭示,复杂图文组合引发更强跨模态交互,降低模型生成拒绝回应的置信度。这些发现凸显信息过载是真实部署中亟需应对的新兴安全风险,亟需加强多模态越狱防御。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, their dual-modal attack surfaces make them vulnerable to jailbreak attacks. Existing LVLM jailbreaks rely on simple designs, e.g., short text and out-of-distribution images. Nevertheless, recent advancements in both large language model backbones and multimodal mechanisms undermine these attacks, particularly their transferability among model architectures. To overcome this limitation, we propose a novel information overloading method that is equipped with both extensive text and multi-dimensional image attacks. These components are arranged in recursion-based image-typography layouts to exponentially increase multimodal information complexity. This overloading approach amplifies the cross-modal processing required, which undermines the safety alignment in LVLMs. Extensive experiments on both open-sourced and commercial LVLMs establish our method as a new state-of-the-art LVLM jailbreak attack. On open-source models, our method achieves an average ASR of 88.6%; on commercial LVLMs, it reaches an average ASR of 84.0%, exceeding the best baseline by 48.7%. Moreover, our prompts optimized on open-source surrogate models transfer effectively across model families. Beyond empirical results, we probe the safety-critical information flows within victim LVLMs. Our observations reveal that complex image-typography compositions induce intensified cross-modal processing and reduce the model's certainty in generating refusal responses. Together, these findings highlight information overloading as a practical and emerging safety risk for real-world LVLM deployments, underscoring the need for stronger defenses against complex multimodal jailbreak inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。