通过分布对齐减少大模型安全对齐带来的推理能力损失。
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
- 用目标模型内部分布重构安全数据,缩小分布差距。
- 推理准确率提升30.2%(DirectRefusal)和21.2%(R1-ACT)。
- 仅需10个样本即可激活有效拒答行为,揭示安全机制本质。
安全对齐会引发安全税,损害大型推理模型(LRM)的通用推理能力。现有用于安全对齐的数据集通常由外部LRM或人工标注者提炼而来,但其推理轨迹与目标模型存在分布差异,我们推测这是导致推理能力显著下降的原因。为此,提出DGR方法,将非目标分布的安全推理数据转换为与目标模型内部分布一致的数据。实验表明:i)DGR在保持安全性能的同时显著缓解安全税,相比原始SFT,在DirectRefusal上平均推理准确率提升30.2%,在R1-ACT上提升21.2%;ii)推理退化程度与分布偏移程度正相关,说明弥合分布差距是保留能力的核心。此外发现,安全对齐主要起激活隐含知识的作用,仅需10个样本即可触发有效拒答行为。这些结果强调了分布一致性的重要性,并揭示了推理模型中安全机制的激活原理。
原文摘要 · Abstract (English)
Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually constructed by distilling safety reasoning traces and answers from an external LRM or human labeler. However, such reasoning traces and answers exhibit a distributional gap with the target LRM that needs alignment, and we conjecture such distributional gap is the culprit leading to significant degradation of reasoning ability of the target LRM. Driven by this hypothesis, we propose a safety alignment dataset construction method, dubbed DGR. DGR transforms and refines an existing out-of-distributional safety reasoning dataset to be aligned with the target's LLM inner distribution. Experimental results demonstrate that i) DGR effectively mitigates the safety tax while maintaining safety performance across all baselines, i.e., achieving \textbf{+30.2\%} on DirectRefusal and \textbf{+21.2\%} on R1-ACT improvement in average reasoning accuracy compared to Vanilla SFT; ii) the degree of reasoning degradation correlates with the extent of distribution shift, suggesting that bridging this gap is central to preserving capabilities. Furthermore, we find that safety alignment in LRMs may primarily function as a mechanism to activate latent knowledge, as a mere \textbf{10} samples are sufficient for activating effective refusal behaviors. These findings not only emphasize the importance of distributional consistency but also provide insights into the activation mechanism of safety in reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。