arXiv:2504.08813cs.LGcs.AI2025-04被引 33

首次系统分析多模态大模型推理安全,发现推理能力越强,越容易被攻破。

SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models

  • 对比基础模型与推理增强模型,发现推理能力显著降低安全对齐
  • 攻击成功率高出37.44%,特定场景攻击率甚至飙升25倍
  • 模型能自发纠正16.9%的危险推理步骤,暗示潜在自愈机制

多模态大推理模型(MLRMs)在众多应用中迅速发展,但其安全性尚未得到充分研究。本文通过大规模实证分析,首次系统比较了具备推理能力的MLRMs与其基础多模态语言模型(MLLMs)的安全性差异。实验揭示三大关键发现:(1)推理税现象:获得推理能力后,模型安全对齐性能急剧下降,对抗攻击下MLRMs的越狱成功率比基线模型高37.44%;(2)安全盲区:尽管安全退化普遍存在,某些场景(如非法活动)的攻击率高达平均值的25倍,远超平均3.4倍增幅,且跨模型、跨数据集表现一致;(3)涌现式自我修正:尽管推理与回答高度耦合,仍有16.9%的越狱推理步骤被安全回复覆盖,提示内在防护机制存在。为推动研究,我们开源OpenSafeMLRM,首个支持主流模型、数据集和越狱方法的统一评估工具包。本工作呼吁立即加强推理增强型AI的安全防护,确保其技术潜力与伦理规范同步。

原文摘要 · Abstract (English)

The rapid advancement of multi-modal large reasoning models (MLRMs) -- enhanced versions of multimodal language models (MLLMs) equipped with reasoning capabilities -- has revolutionized diverse applications. However, their safety implications remain underexplored. While prior work has exposed critical vulnerabilities in unimodal reasoning models, MLRMs introduce distinct risks from cross-modal reasoning pathways. This work presents the first systematic safety analysis of MLRMs through large-scale empirical studies comparing MLRMs with their base MLLMs. Our experiments reveal three critical findings: (1) The Reasoning Tax: Acquiring reasoning capabilities catastrophically degrades inherited safety alignment. MLRMs exhibit 37.44% higher jailbreaking success rates than base MLLMs under adversarial attacks. (2) Safety Blind Spots: While safety degradation is pervasive, certain scenarios (e.g., Illegal Activity) suffer 25 times higher attack rates -- far exceeding the average 3.4 times increase, revealing scenario-specific vulnerabilities with alarming cross-model and datasets consistency. (3) Emergent Self-Correction: Despite tight reasoning-answer safety coupling, MLRMs demonstrate nascent self-correction -- 16.9% of jailbroken reasoning steps are overridden by safe answers, hinting at intrinsic safeguards. These findings underscore the urgency of scenario-aware safety auditing and mechanisms to amplify MLRMs' self-correction potential. To catalyze research, we open-source OpenSafeMLRM, the first toolkit for MLRM safety evaluation, providing unified interface for mainstream models, datasets, and jailbreaking methods. Our work calls for immediate efforts to harden reasoning-augmented AI, ensuring its transformative potential aligns with ethical safeguards.

多模态安全评测大模型推理越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。