arXiv:2503.00555cs.CRcs.AI2025-03被引 133

安全对齐让大模型更安全,但会牺牲推理能力。

Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

  • 在大推理模型上实施安全对齐,恢复其安全性。
  • 安全对齐使推理能力下降,存在权衡关系。
  • 适合关注模型安全与性能平衡的研究者。

安全对齐是大型语言模型(LLM)正式部署前的重要步骤。尽管对LLM的安全对齐已有广泛研究,但具备更强推理能力的大推理模型(LRM)相关研究仍存在较大空白。本文系统考察了一种简化的生成安全对齐LRM的流程。通过评估多种LRM,得出两个主要发现:(i) 可在LRM上实施安全对齐以恢复其安全能力;(ii) 安全对齐会导致LRM推理能力下降。这两个发现表明,在顺序式LRM生成流程中,推理能力与安全能力之间存在权衡关系。我们称这一现象为‘安全税’,它应为未来针对LRM的安全研究提供启示。作为副产品,我们构建了一个名为DirectRefusal的数据集,可能可作为安全对齐的替代数据集。代码已开源于https://github.com/git-disl/Safety-Tax。

原文摘要 · Abstract (English)

Safety alignment is an important procedure before the official deployment of a Large Language Model (LLM). While safety alignment has been extensively studied for LLM, there is still a large research gap for Large Reasoning Models (LRMs) that equip with improved reasoning capability. We in this paper systematically examine a simplified pipeline for producing safety aligned LRMs. With our evaluation of various LRMs, we deliver two main findings: i) Safety alignment can be done upon the LRM to restore its safety capability. ii) Safety alignment leads to a degradation of the reasoning capability of LRMs. The two findings show that there exists a trade-off between reasoning and safety capability with the sequential LRM production pipeline. The discovered trade-off, which we name Safety Tax, should shed light on future endeavors of safety research on LRMs. As a by-product, we curate a dataset called DirectRefusal, which might serve as an alternative dataset for safety alignment. Our source code is available at https://github.com/git-disl/Safety-Tax.

大模型安全推理能力安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。