arXiv:2510.04196cs.AIcs.LG2025-10

让大模型安全与能力共成长,防滥用还少拒答

COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

  • 用多任务多目标强化学习联合训练,避免安全与能力互相压制
  • 在多模态攻击下拒答率降37%,指令遵循能力提升12%
  • 适配不同模型架构,适合需高安全性的实际应用

大型多模态推理模型(LMRMs)正迈向真实应用,必须兼顾实用性和安全性。多模态场景下,图像与文本结合易绕过安全限制,单一目标训练会导致策略漂移,引发对无害输入过度拒绝或对风险内容不当响应。本文提出COSMO-RL框架,通过多模态、多任务、多目标强化学习联合训练推理型LMRMs,发布模型COSMO-R1。该方法旨在让安全与能力在稳定流程中共同提升,而非对齐阶段相互竞争。实验表明,COSMO-R1在保持甚至提升多模态推理与指令遵循能力的同时显著增强安全性,对多模态越狱攻击的鲁棒性更强,且减少不必要的拒绝行为。该框架可迁移至不同主干网络,均取得一致性能提升。消融实验验证了设计有效性,为实现安全与通用能力协同进化提供了简洁路径。

原文摘要 · Abstract (English)

Large Multimodal Reasoning Models (LMRMs) are moving into real applications, where they must be both useful and safe. Safety is especially challenging in multimodal settings: images and text can be combined to bypass guardrails, and single objective training can cause policy drift that yields over-refusal on benign inputs or unsafe compliance on risky ones. We present COSMO-RL, a mixed reinforcement learning framework that trains reasoning oriented LMRMs under multimodal, multitask, and multiobjective signals, and we release the resulting model, COSMO-R1. Our approach aims to let safety and capability grow together in one stable pipeline rather than competing during alignment. In experiments, COSMO-R1 improves safety while maintaining-and often improving multimodal reasoning and instruction following, shows stronger robustness to multimodal jailbreaks, and reduces unnecessary refusals. The framework also transfers across backbones with consistent gains. Ablations support the design choices, indicating a simple path to advancing safety and general capability together in LMRMs.

多模态安全对齐强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。