arXiv:2512.01848cs.CL2025-12被引 1

用强化学习提升大模型推理安全性,兼顾安全与逻辑能力。

Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability

  • 采用强化学习优化推理过程,动态调整行为策略。
  • 在多个模型上实现更稳定的安全性提升,推理能力不下降。
  • 适合关注大模型安全与可靠推理的研究者。

大推理模型(LRMs)通过生成显式的思维链(CoT)显著提升数学与逻辑问题求解能力。然而,这一显式推理过程也引入新安全风险:即使最终答案看似无害,中间推理轨迹仍可能包含不当行为。现有安全对齐方法主要依赖基于安全长思维链数据集的监督微调(SFT),但实验发现其安全提升不稳定、削弱推理能力且跨模型泛化差。为此,本文探索强化学习(RL)作为互补优化框架。相比SFT,RL通过奖励反馈直接优化模型策略,实现更自适应、稳定的对齐。多模型家族与基准测试显示,RL在保持推理能力的同时,带来更强且一致的安全性提升。进一步分析揭示,RL抑制了不安全的探索性推理,同时保留反思深度,使推理过程更安全可靠。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) extend large language models by generating explicit chain-of-thought (CoT) reasoning, significantly improving mathematical and logical problem solving. However, this explicit reasoning process also introduces new safety risks, as unsafe behaviors often emerge within intermediate reasoning trajectories, even when final answers appear harmless. Existing safety alignment approaches primarily rely on supervised fine-tuning (SFT) over safety-oriented long CoT datasets. While intuitive, we find that SFT produces inconsistent safety improvements, degrades reasoning ability, and generalizes poorly across model families. These limitations suggest that purely supervised approaches are insufficient for robust safety alignment in LRMs. To address this, we investigate reinforcement learning (RL) as a complementary optimization framework for LRM safety training. Unlike SFT, RL directly optimizes model policies with reward feedback, enabling more adaptive and stable alignment. Extensive experiments across multiple model families and benchmarks show that RL achieves stronger and more consistent safety gains while maintaining reasoning competence. Further analysis of reflection dynamics and token-level entropy reveals that RL suppresses unsafe exploratory reasoning while preserving reflective depth, leading to safer and more reliable reasoning processes.

大模型安全强化学习推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。