推理模型的安全行为其实早被第一句话决定,思考过程多是伪装的。
Do Thinking Tokens Help with Safety?

- 用首词隐藏表示就能预测模型是否拒绝或同意,准确率超88%
- 90%以上思考文本看似在权衡,实际结果早已锁定
- 现有安全策略反而让模型更倾向于过度拒绝,抑制真实思考
当前推理模型依赖思考标记以在基准测试中超越指令微调模型。普遍认为这种更‘审慎’的模式能提升对齐与安全性,为模型提供空间判断回应是否违反安全原则。我们发现这一直觉并不总是成立。在涵盖GPT-OSS、Qwen、Olmo和Phi系列的前沿开源推理模型中,最终拒绝或同意的结果可通过首词隐藏表示的训练头提前预测,AUROC达0.84–0.95,平衡准确率约88%。思考过程实为前缀补全而非审慎修订,最终结果在约20%思考阶段后基本固定,尽管文本层面呈现74%的表观权衡。现有基于推理时或训练的安全干预虽旨在促进审慎思考,却大多导致过度拒绝,同时压制本已稀少的审慎信号。结果表明,当前推理模型的安全行为远不如预期般审慎,亟需能真正诱导安全审慎的新方法。
原文摘要 · Abstract (English)
Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also generally believed that this more "deliberative" mode should improve alignment and safety, by providing the model a safe space to consider whether its planned answer to a request violates its safety principles. We present evidence that this intuition is not always correct. Across frontier open-weight reasoning models spanning GPT-OSS, Qwen, Olmo, and Phi families, we find that the eventual refusal/compliance outcome is already strongly predictable via a trained head on the first token's hidden representation ($0.84$-$0.95$ AUROC and $\sim88\%$ balanced accuracy for predicting refusal/compliance) before any visible thinking. The thinking process turns out to be more akin to prefix completion than to deliberative revision, with the final outcome rarely changing after the first $\sim20\%$ of thinking, despite giving the appearance of deliberation at the text level ($\sim74\%$ of text-level deliberations occur when the response distribution is already locked to one refusal/compliance side). We also find that existing inference-time and training-based safety interventions, despite being motivated by the goal of inducing deliberation, largely shift model behavior toward over-refusal while suppressing already-scarce deliberation signals. Our results suggest that safety behavior in current reasoning models is much less deliberative than commonly assumed, and highlight the need for methods that induce real safety deliberation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。