arXiv:2503.17682cs.LGcs.AI2025-03NeurIPS被引 25

首个多模态安全对齐框架,兼顾模型有用性与安全性

Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

  • 基于双偏好信号设计安全约束优化,实现多模态模型的安全微调
  • 安全性能提升34.2%,有用性提升34.3%,在五轮过滤后平均安全增益达40.9%
  • 适用于需要高安全性的通用多模态AI助手开发

多模态大语言模型是构建通用AI助手的关键,但其安全风险日益突出。如何确保多模态大模型的安全对齐以避免不当行为?进一步地,如何在保持能力的同时满足安全约束?该问题可形式化为极小极大优化问题。然而,现有数据集尚未将单一偏好信号解耦为明确的安全约束,阻碍了系统性研究。此外,这类约束能否有效融入多模态模型的优化过程仍属开放问题。本文提出首个多模态安全对齐框架Safe RLHF-V,包含三部分:(I) BeaverTails-V,首个开源双偏好数据集,标注帮助性与安全性,并附有三级安全标签(轻微、中度、严重);(II) Beaver-Guard-V,多级防护系统,通过五轮过滤与重生成,使预训练模型整体安全性平均提升40.9%;(III) 基于双偏好信号,首次探索多模态安全约束下的优化。实验表明,Safe RLHF-V能有效提升模型的帮助性与安全性,分别提升34.2%和34.3%。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: $\mathbf{(I)}$ BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); $\mathbf{(II)}$ Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model's overall safety by an average of 40.9%. $\mathbf{(III)}$ Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%.

多模态安全对齐强化学习人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。