通过视觉推理逐步诱导多模态大模型输出有害内容
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
- 将有害文本拆解为连续子图像,逐步引导模型暴露恶意意图
- 在GPT-4o和Claude-4.5-Sonnet上攻击成功率超越现有方法
- 适合研究多模态安全与对抗攻击的学者参考
多模态大语言模型(MLLMs)因强大的跨模态理解与生成能力被广泛应用,但更多模态也带来了新的越狱攻击漏洞,可能诱导其输出有害内容。尽管MLLMs具备强推理能力,现有越狱攻击主要聚焦于文本模态,而视觉模态的安全风险被严重忽视。为此,我们提出视觉推理序列攻击(VRSA),通过将原始有害文本分解为多个语义关联的子图像序列,逐步诱导MLLM外化并聚合完整恶意意图。为提升图像序列场景合理性,提出自适应场景优化(Adaptive Scene Refinement);为保障生成图像的语义连贯性,提出语义一致性补全(Semantic Coherent Completion),迭代融合上下文信息重写子文本;同时引入文本-图像一致性对齐(Text-Image Consistency Alignment)维持语义一致。实验表明,VRSA在开源与闭源MLLM(如GPT-4o、Claude-4.5-Sonnet)上均显著优于当前最优攻击方法。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are widely used in various fields due to their powerful cross-modal comprehension and generation capabilities. However, more modalities bring more vulnerabilities to being utilized for jailbreak attacks, which induces MLLMs to output harmful content. Due to the strong reasoning ability of MLLMs, previous jailbreak attacks try to explore reasoning safety risk in text modal, while similar threats have been largely overlooked in the visual modal. To fully evaluate potential safety risks in the visual reasoning task, we propose Visual Reasoning Sequential Attack (VRSA), which induces MLLMs to gradually externalize and aggregate complete harmful intent by decomposing the original harmful text into several sequentially related sub-images. In particular, to enhance the rationality of the scene in the image sequence, we propose Adaptive Scene Refinement to optimize the scene most relevant to the original harmful query. To ensure the semantic continuity of the generated image, we propose Semantic Coherent Completion to iteratively rewrite each sub-text combined with contextual information in this scene. In addition, we propose Text-Image Consistency Alignment to keep the semantical consistency. A series of experiments demonstrates that the VRSA can achieve a higher attack success rate compared with the state-of-the-art jailbreak attack methods on both the open-source and closed-source MLLMs such as GPT-4o and Claude-4.5-Sonnet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。