解决多模态大模型因推理链过长导致的安全隐患问题
When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- 构建可解释的跨模态安全推理数据集SSUI
- 训练后模型在安全基准上显著优于主流模型
- 适合关注多模态模型安全对齐的研究者
多模态大语言模型(MLLMs)存在隐式推理风险:看似无害的单模态输入在组合后可能生成有害输出。我们发现该漏洞源于模型在长链推理中难以保持安全对齐。为此,提出首个面向跨模态安全挑战的可解释推理路径数据集SSUI,并设计基于该数据集的安全感知推理链优化框架SRPO,使模型内部推理过程与人类安全价值观对齐。实验表明,经SRPO训练的模型在关键安全基准(包括自建的推理路径基准RSBench)上达到当前最优表现,显著优于开源及顶尖商业MLLMs。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are susceptible to the implicit reasoning risk, wherein innocuous unimodal inputs synergistically assemble into risky multimodal data that produce harmful outputs. We attribute this vulnerability to the difficulty of MLLMs maintaining safety alignment through long-chain reasoning. To address this issue, we introduce Safe-Semantics-but-Unsafe-Interpretation (SSUI), the first dataset featuring interpretable reasoning paths tailored for such a cross-modal challenge. A novel training framework, Safety-aware Reasoning Path Optimization (SRPO), is also designed based on the SSUI dataset to align the MLLM's internal reasoning process with human safety values. Experimental results show that our SRPO-trained models achieve state-of-the-art results on key safety benchmarks, including the proposed Reasoning Path Benchmark (RSBench), significantly outperforming both open-source and top-tier commercial MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。