arXiv:2508.00324cs.AIcs.CL2025-08被引 4

让大模型主动调用安全知识,只需千条数据和半天训练

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge

  • 通过结构化推理触发模型内置的安全知识
  • 仅用1000样本90分钟训练,安全性能显著提升
  • 适合希望快速部署安全增强的大模型应用者

尽管大型推理模型在复杂任务上表现出色,但近期研究发现这些模型常执行有害用户指令,引发严重安全问题。本文探究其根源,发现模型本身已具备充足安全知识,但在推理过程中未能激活。基于此,提出R1-Act——一种简单高效的后训练方法,通过结构化推理过程显式触发安全知识。该方法在保持推理能力的同时显著提升安全性,优于以往对齐方法。特别地,仅需1000个训练样本和单块RTX A6000 GPU上90分钟训练即可完成。在多个大型推理模型架构与规模下进行的广泛实验表明,该方法具备强鲁棒性、可扩展性和实际效率。

原文摘要 · Abstract (English)

Although large reasoning models (LRMs) have demonstrated impressive capabilities on complex tasks, recent studies reveal that these models frequently fulfill harmful user instructions, raising significant safety concerns. In this paper, we investigate the underlying cause of LRM safety risks and find that models already possess sufficient safety knowledge but fail to activate it during reasoning. Based on this insight, we propose R1-Act, a simple and efficient post-training method that explicitly triggers safety knowledge through a structured reasoning process. R1-Act achieves strong safety improvements while preserving reasoning performance, outperforming prior alignment methods. Notably, it requires only 1,000 training examples and 90 minutes of training on a single RTX A6000 GPU. Extensive experiments across multiple LRM backbones and sizes demonstrate the robustness, scalability, and practical efficiency of our approach.

模型安全推理增强后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。