用漫画式连环画绕过多模态大模型安全机制,成功率超83%。
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- 将恶意请求拆解为无害视觉叙事元素,通过扩散模型生成图像序列
- 在多个安全基准上实现83.5%平均攻击成功率,比现有方法高46%
- 揭示多模态安全机制漏洞,适合研究模型防御与对抗攻击者
多模态大语言模型(MLLMs)虽具强大能力,但仍易受跨模态漏洞引发的越狱攻击。本文提出一种新方法,利用漫画风格的连续视觉叙事绕过顶尖MLLM的安全对齐机制。该方法借助辅助大模型将恶意查询分解为视觉无害的叙事元素,通过扩散模型生成图像序列,并利用模型对叙事连贯性的依赖诱导出有害输出。在来自权威安全基准的有害文本查询上进行的大量实验表明,该方法平均攻击成功率达83.5%,较先前最优方法提升46%。相比现有视觉越狱方法,本策略在多种有害内容类别中均表现出更优效果。我们进一步分析攻击模式,揭示多模态安全机制中的关键脆弱因素,并评估当前防御策略在应对叙事驱动攻击时的局限性,暴露出现有防护体系的重大短板。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) exhibit remarkable capabilities but remain susceptible to jailbreak attacks exploiting cross-modal vulnerabilities. In this work, we introduce a novel method that leverages sequential comic-style visual narratives to circumvent safety alignments in state-of-the-art MLLMs. Our method decomposes malicious queries into visually innocuous storytelling elements using an auxiliary LLM, generates corresponding image sequences through diffusion models, and exploits the models' reliance on narrative coherence to elicit harmful outputs. Extensive experiments on harmful textual queries from established safety benchmarks show that our approach achieves an average attack success rate of 83.5\%, surpassing prior state-of-the-art by 46\%. Compared with existing visual jailbreak methods, our sequential narrative strategy demonstrates superior effectiveness across diverse categories of harmful content. We further analyze attack patterns, uncover key vulnerability factors in multimodal safety mechanisms, and evaluate the limitations of current defense strategies against narrative-driven attacks, revealing significant gaps in existing protections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。