通过跨模态纠缠攻击,突破视觉语言模型的安全限制。
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
- 将有害指令拆解为多跳链式任务,增强攻击复杂度。
- 在图像中嵌入可可视化实体,构建跨模态推理路径。
- 适合用于测试视觉语言模型的持续安全防御能力。
具备多模态推理能力的视觉语言模型(VLMs)是高价值攻击目标,因其能处理复杂的多模态有害任务。当前主流黑盒越狱攻击通过在不同模态间分散恶意线索来转移模型注意力,绕过安全对齐机制。然而,这些攻击依赖简单固定的图文组合,缺乏攻击复杂度的可扩展性,难以有效检验VLMs持续演进的推理能力。本文提出 extbf{CrossTALK}(跨模态纠缠攻击),一种可扩展的方法,通过跨模态信息纠缠突破VLMs预训练的安全对齐模式。具体包括:知识可扩展重构,将有害任务扩展为多跳链式指令;跨模态线索纠缠,将可可视化实体迁移至图像以建立多模态推理关联;跨模态场景嵌套,利用多模态上下文指令引导模型生成详细有害输出。实验表明,所提方法在攻击成功率上达到当前最优水平。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) with multimodal reasoning capabilities are high-value attack targets, given their potential for handling complex multimodal harmful tasks. Mainstream black-box jailbreak attacks on VLMs work by distributing malicious clues across modalities to disperse model attention and bypass safety alignment mechanisms. However, these adversarial attacks rely on simple and fixed image-text combinations that lack attack complexity scalability, limiting their effectiveness for red-teaming VLMs' continuously evolving reasoning capabilities. We propose \textbf{CrossTALK} (\textbf{\underline{Cross}}-modal en\textbf{\underline{TA}}ng\textbf{\underline{L}}ement attac\textbf{\underline{K}}), which is a scalable approach that extends and entangles information clues across modalities to exceed VLMs' trained and generalized safety alignment patterns for jailbreak. Specifically, {knowledge-scalable reframing} extends harmful tasks into multi-hop chain instructions, {cross-modal clue entangling} migrates visualizable entities into images to build multimodal reasoning links, and {cross-modal scenario nesting} uses multimodal contextual instructions to steer VLMs toward detailed harmful outputs. Experiments show our COMET achieves state-of-the-art attack success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。