研究推理模型何时拒绝有害请求,发现初始推理句决定拒绝行为
Where Do Reasoning Models Refuse?
- 通过固定推理过程,发现推理链因果影响拒绝结果
- 精炼模型中,推理开头句差异可完全决定是否拒绝
- 提取拒绝方向并消融测试,发现提升安全但削弱通用能力
无思维链(CoT)的聊天模型必须在生成首个输出标记前决定是否拒绝有害请求。而推理模型在最终输出前会生成较长的推理链,引发一个自然问题:拒绝决策究竟发生在哪个阶段?我们针对四款开源推理模型展开研究。首先表明,思维链对拒绝结果具有因果影响;固定特定推理路径可显著降低模型最终拒绝或遵从的波动性。深入分析推理链发现,在精炼模型中,推理开头句的细微差异即可完全决定拒绝决策,且此类模式可在来自同一教师模型的其他模型间迁移。最后,我们从模型激活中提取线性拒绝方向,并发现消融该方向会增加有害行为的顺从率,但效果不如非推理模型显著,且会对通用能力造成不可忽视的下降。
原文摘要 · Abstract (English)
Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response token. Reasoning models, by contrast, produce extended chains of thought before their final output, raising a natural question: where in this process does the decision to refuse occur? We investigate this across four open-source reasoning models. We first show that the CoT causally influences refusal outcomes; fixing a specific reasoning trace substantially reduces variance in whether the model ultimately refuses or complies. Zooming into the reasoning trace, we find that in distilled models, subtle differences in the opening sentence of the CoT can fully determine the model's refusal decision, and that these patterns transfer across models distilled from the same teacher. Finally, we extract linear refusal directions from model activations and show that ablating them increases harmful compliance, though less reliably than the same technique achieves on non-reasoning models, and with non-negligible degradation to general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。