推理模型虽提升解题能力,却常丧失安全与伦理对齐性。
Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models

- 通过监督微调、强化学习和蒸馏生成推理模型
- 多维度测试显示毒性、偏见、隐私泄露等问题上升
- 建议评估推理模型时必须包含可信度指标
指令微调的大语言模型通过后训练转为推理模型以提升多步任务表现,但该过程通常只优化推理准确性,未显式保持原始模型的安全拒绝、去偏和隐私保护等对齐行为。我们通过可信度审计发现,这种转换默认不保留对齐性。系统比较了三种方法(监督微调、基于强化学习的后训练、蒸馏)生成的推理模型与对应指令微调基线,在六个可信度维度(安全、毒性、刻板印象与偏见、机器伦理、隐私、分布外鲁棒性)上的表现。结果表明,推理模型在推理基准上表现更好,但出现对齐退化:毒性升高、刻板印象加剧、拒绝行为校准失准、上下文隐私泄露。这些退化与指令微调基线存在行为漂移(以KL散度衡量)。结论强调,评估推理模型时必须报告可信度指标,不能仅关注推理能力提升。
原文摘要 · Abstract (English)
Instruction-tuned LLMs are increasingly converted into reasoning models through post-training to improve multi-step task performance. This conversion is usually optimized for reasoning accuracy, without explicitly preserving the alignment behavior of the instruction-tuned model, such as safe refusal, bias avoidance, and privacy protection. We ask: does this conversion preserve alignment? We study this question through a trustworthiness audit and find that it is not behavior-preserving by default. For a systematic analysis, we compare reasoning models produced via supervised fine-tuning, RL-based post-training, and distillation against matched instruction-tuned baselines across six trustworthiness dimensions: safety, toxicity, stereotyping and bias, machine ethics, privacy, and out-of-distribution robustness. We observe that reasoning models often improve on reasoning benchmarks but exhibit alignment regressions, including increased toxicity, amplified stereotyping, miscalibrated refusal, and contextual privacy leakage. These regressions are consistent with behavioral drift from the instruction-tuned baseline, measured by KL divergence. Overall, our results point to the broader conclusion that trustworthiness metrics are essential for evaluating reasoning models and should be reported alongside gains in reasoning capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。