arXiv:2505.19684cs.CV2025-05EMNLP被引 12

视觉推理越强,模型越易被攻破,新攻击方法可突破主流多模态大模型安全防线。

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models

  • 利用视觉推理链设计分阶段诱导攻击,精准控制有害输出。
  • 在三款主流闭源模型上攻击成功率超56%,最高达76.48%。
  • 揭示视觉推理能力与安全风险间的根本矛盾,适合安全研究者参考。

多模态大语言模型(MLRMs)通过结合强化学习与思维链(CoT)监督,提升了复杂的视觉推理能力。然而,这种增强的推理能力也带来了未被充分探索的安全风险。本文系统研究了先进视觉推理对MLRMs安全性的影响,发现一个核心矛盾:视觉推理能力越强,模型越容易受到越狱攻击。为此,我们提出VisCRA(视觉链式推理攻击),一种利用视觉推理链绕过安全机制的新框架。VisCRA结合目标视觉注意力掩码与两阶段推理诱导策略,实现对有害输出的精确控制。大量实验表明,VisCRA在主流闭源模型上表现显著:在Gemini 2.0 Flash Thinking上成功率达76.48%,QvQ-Max为68.56%,GPT-4o为56.60%。研究揭示:赋予模型强大视觉推理能力的同时,也可能成为攻击入口,带来严重安全隐患。

原文摘要 · Abstract (English)

The emergence of Multimodal Large Language Models (MLRMs) has enabled sophisticated visual reasoning capabilities by integrating reinforcement learning and Chain-of-Thought (CoT) supervision. However, while these enhanced reasoning capabilities improve performance, they also introduce new and underexplored safety risks. In this work, we systematically investigate the security implications of advanced visual reasoning in MLRMs. Our analysis reveals a fundamental trade-off: as visual reasoning improves, models become more vulnerable to jailbreak attacks. Motivated by this critical finding, we introduce VisCRA (Visual Chain Reasoning Attack), a novel jailbreak framework that exploits the visual reasoning chains to bypass safety mechanisms. VisCRA combines targeted visual attention masking with a two-stage reasoning induction strategy to precisely control harmful outputs. Extensive experiments demonstrate VisCRA's significant effectiveness, achieving high attack success rates on leading closed-source MLRMs: 76.48% on Gemini 2.0 Flash Thinking, 68.56% on QvQ-Max, and 56.60% on GPT-4o. Our findings highlight a critical insight: the very capability that empowers MLRMs -- their visual reasoning -- can also serve as an attack vector, posing significant security risks.

越狱攻击多模态模型安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。