arXiv:2507.18576cs.AIcs.CL2025-07被引 2

安全与智能共进化,让大模型自己学会守规矩。

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law

  • 用渐进式安全强化学习,让模型自发产生安全意识。
  • 安全性能提升46.54%,超越GPT-4.1和Claude Opus 4。
  • 适合追求高安全性的通用AI开发与部署场景。

我们提出SafeWork-R1,一种多模态推理模型,实现了能力与安全的共进化。该模型基于我们提出的SafeLadder框架,采用大规模、渐进式、以安全为导向的强化学习后训练,并配备多原则验证器。与传统RLHF仅学习人类偏好不同,SafeLadder使SafeWork-R1具备内在安全推理与自我反思能力,产生安全“顿悟”时刻。值得注意的是,SafeWork-R1在安全基准测试中相较其基线模型Qwen2.5-VL-72B平均提升46.54%,且不牺牲通用能力,安全表现优于GPT-4.1和Claude Opus 4等领先专有模型。为进一步增强可靠性,我们引入两种推理时干预方法与分步验证的思辨搜索机制。此外,我们还构建了SafeWork-R1-InternVL3-78B、SafeWork-R1-DeepSeek-70B和SafeWork-R1-Qwen2.5VL-7B,均表明安全与能力可协同进化,凸显框架在构建可靠可信通用AI中的普适性。

原文摘要 · Abstract (English)

We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framework, which incorporates large-scale, progressive, safety-oriented reinforcement learning post-training, supported by a suite of multi-principled verifiers. Unlike previous alignment methods such as RLHF that simply learn human preferences, SafeLadder enables SafeWork-R1 to develop intrinsic safety reasoning and self-reflection abilities, giving rise to safety `aha' moments. Notably, SafeWork-R1 achieves an average improvement of $46.54\%$ over its base model Qwen2.5-VL-72B on safety-related benchmarks without compromising general capabilities, and delivers state-of-the-art safety performance compared to leading proprietary models such as GPT-4.1 and Claude Opus 4. To further bolster its reliability, we implement two distinct inference-time intervention methods and a deliberative search mechanism, enforcing step-level verification. Finally, we further develop SafeWork-R1-InternVL3-78B, SafeWork-R1-DeepSeek-70B, and SafeWork-R1-Qwen2.5VL-7B. All resulting models demonstrate that safety and capability can co-evolve synergistically, highlighting the generalizability of our framework in building robust, reliable, and trustworthy general-purpose AI.

安全对齐多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。