arXiv:2505.06538cs.CL2025-05EMNLP被引 21

发现多模态大模型安全能力会随推理变长而下降,提出用推理过程主动防危险。

Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model

  • 利用模型自身推理过程识别潜在危险意图
  • 在5个基准上测试11个模型,发现多数存在安全退化现象
  • 构建安全导向推理数据集,可有效提升模型安全性

多模态大推理模型(MLRMs)发展迅速,但其安全性和可靠性仍存隐患。本文对11个MLRMs在5个基准上进行了系统性安全评估,揭示了多数先进模型普遍存在安全性能下降现象。分析显示,不同基准表现差异显著:越狱鲁棒性基准中安全退化明显,而安全意识类基准则退化较弱;尤其在某些需要长推理链的场景中,推理过程反而提升了安全性。这表明可借助模型内在推理能力主动检测不安全意图。为此,我们构建了一个包含安全导向推理过程的多模态微调数据集。实验表明,使用该数据集对现有MLRMs进行微调,能有效提升其在越狱鲁棒性和安全意识类基准上的表现。研究为构建更安全的MLRMs提供了新思路。数据集已开源:https://github.com/xinyuelou/Think-in-Safety。

原文摘要 · Abstract (English)

The rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration. To address this gap, we conduct a comprehensive and systematic safety evaluation of 11 MLRMs across 5 benchmarks and unveil prevalent safety degradation phenomena in most advanced models. Moreover, our analysis reveals distinct safety patterns across different benchmarks: significant safety degradation is observed across jailbreak robustness benchmarks, whereas safety-awareness benchmarks demonstrate less pronounced degradation. In particular, the long thought process in some scenarios even enhances safety performance. Therefore, it is a potential approach to address safety issues in MLRMs by leveraging the intrinsic reasoning capabilities of the model to detect unsafe intent. To operationalize this insight, we construct a multimodal tuning dataset that incorporates a safety-oriented thought process. Experimental results from fine-tuning existing MLRMs with this dataset effectively enhances the safety on both jailbreak robustness and safety-awareness benchmarks. This study provides a new perspective for developing safe MLRMs. Our dataset is available at https://github.com/xinyuelou/Think-in-Safety.

多模态安全对齐推理增强模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。