改变推理结构可显著提升大模型安全性
Reasoning Structure Matters for Safety Alignment of Reasoning Models

- 通过修改推理结构实现安全对齐
- 仅用1000条数据微调即有效
- 无需强化学习,适配多场景
大型推理模型(LRMs)在复杂推理任务中表现优异,但常对恶意用户查询生成有害回应。本文揭示其安全风险源于推理结构本身。基于此,提出AltTrain:一种仅需1000条标注数据的监督微调方法,通过显式改变推理结构实现安全对齐。该方法无需复杂强化学习或奖励设计,具有高度实用性与泛化性。实验表明,其在多种模型架构与规模下均能实现强安全对齐,并在推理、问答、摘要及多语言场景中表现稳健。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose AltTrain, a simple yet effective post training method that explicitly alters the reasoning structure of LRMs. AltTrain is both practical and generalizable, requiring no complex reinforcement learning (RL) training or reward design, only supervised finetuning (SFT) with a lightweight 1K training examples. Experiments across LRM backbones and model sizes demonstrate strong safety alignment, along with robust generalization across reasoning, QA, summarization, and multilingual setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。