通过规则奖励提升多模态模型的安全推理能力
SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- 用规则约束的奖励机制增强自奖励策略优化的安全性
- 在SafeTag-VL-3K数据集上显著提升跨模态安全意识
- 适合关注多模态模型安全对齐的研究者与开发者
多模态大语言模型虽具备出色的推理与指令遵循能力,但其扩展的模态空间引入了由复杂图文交互产生的新型组合安全风险。即使输入本身无害,跨模态耦合仍可能生成不安全语义,暴露出当前多模态模型安全感知的脆弱性。尽管已有工作通过引导模型推理潜在风险来提升安全性,但未经监管的推理过程可能破坏对齐;虽然群组相对策略优化(GRPO)可实现无需人工标注的自奖励精炼,却缺乏可验证的推理安全信号。为此,我们提出SafeGRPO——一种将规则化奖励构建融入GRPO的自奖励多模态安全对齐框架,实现可解释且可验证的推理安全优化。基于构建的SafeTag-VL-3K数据集(含显式视觉、文本及联合安全标签),SafeGRPO通过分步引导安全思考,强化结构化推理与行为对齐,在不牺牲通用能力的前提下,显著提升了多模态安全意识、组合鲁棒性与推理稳定性,覆盖多种基准测试。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new compositional safety risks that emerge from complex text-image interactions. Such cross-modal couplings can produce unsafe semantics even when individual inputs are benign, exposing the fragile safety awareness of current MLLMs. While recent works enhance safety by guiding models to reason about potential risks, unregulated reasoning traces may compromise alignment; although Group Relative Policy Optimization (GRPO) offers self-rewarded refinement without human supervision, it lacks verifiable signals for reasoning safety. To address this, we propose SafeGRPO a self-rewarded multimodal safety alignment framework that integrates rule-governed reward construction into GRPO, enabling interpretable and verifiable optimization of reasoning safety. Built upon the constructed SafeTag-VL-3K dataset with explicit visual, textual, and combined safety tags, SafeGRPO performs step-guided safety thinking to enforce structured reasoning and behavior alignment, substantially improving multimodal safety awareness, compositional robustness, and reasoning stability across diverse benchmarks without sacrificing general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。