用可验证奖励训练大模型,既能提升能力又不牺牲安全
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
- 采用可验证奖励机制,通过客观指标优化模型
- 在5个对抗性安全测试中同时提升推理能力与安全性能
- 适合关注大模型安全部署的研究者和开发者
微调大语言模型(LLMs)在下游任务中通常存在安全-能力权衡问题,即提升任务表现会降低安全性,即使在良性数据集上也是如此。这一现象在监督微调(SFT)和基于人类反馈的强化学习(RLHF)中普遍存在。尽管可验证奖励的强化学习(RLVR)作为一种新兴方法,能在客观可测任务上优化模型,其安全影响仍未知。本文首次对RLVR的安全属性进行了系统的理论与实证分析。理论上,我们推导出在KL约束优化下的安全漂移上界,并证明了消除安全退化的条件。实证上,我们在五个对抗性安全基准上进行广泛实验,结果表明RLVR可在提升推理能力的同时维持甚至增强安全防护。全面消融实验考察了优化算法、模型规模和任务领域的影响。研究挑战了安全与能力不可兼得的普遍假设,表明特定训练方法可同时实现二者,为推理型大模型的安全部署提供关键洞见。
原文摘要 · Abstract (English)
Fine-tuning large language models (LLMs) for downstream tasks typically exhibit a fundamental safety-capability tradeoff, where improving task performance degrades safety alignment even on benign datasets. This degradation persists across standard approaches including supervised finetuning (SFT) and reinforcement learning from human feedback (RLHF). While reinforcement learning with verifiable rewards (RLVR) has emerged as a promising alternative that optimizes models on objectively measurable tasks, its safety implications remain unexplored. We present the first comprehensive theoretical and empirical analysis of safety properties in RLVR. Theoretically, we derive upper bounds on safety drift under KL-constrained optimization and prove conditions under which safety degradation is eliminated. Empirically, we conduct extensive experiments across five adversarial safety benchmarks, demonstrating that RLVR can simultaneously enhance reasoning capabilities while maintaining or improving safety guardrails. Our comprehensive ablation studies examine the effects of optimization algorithms, model scale, and task domains. Our findings challenge the prevailing assumption of an inevitable safety capability trade-off, and establish that a specific training methodology can achieve both objectives simultaneously, providing insights for the safe deployment of reasoning-capable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。