arXiv:2604.01202cs.AI2026-04被引 2

推理模型先做决定,再找理由,而非边想边选。

Therefore I am. I Think

  • 用线性探测发现,决策信号在生成前就已编码于激活值中。
  • 扰动决策方向可使推理过程变长,且改变行为(7%-79%)。
  • 模型常为翻转的决策编造合理解释,说明决策早于思考。

我们探讨大语言推理模型做出选择时,是先思考再决定,还是先决定再思考。本文提供证据表明,可检测的早期决策信号会塑造链式思维过程。具体而言,我们发现简单的线性探测器能以极高置信度从生成前的激活值中解码出工具调用决策,甚至在生成任何推理文本之前即可完成。激活值调控进一步支持因果关系:扰动决策方向会导致推理过程显著延长,并在多数例子中改变行为(在不同模型和基准下,变化比例为7%-79%)。行为分析还显示,当调控改变决策时,链式思维过程往往主动为新决策辩护,而非抵抗。这些结果表明,推理模型在开始文本推理之前,可能已预先编码了行动选择。

原文摘要 · Abstract (English)

We consider the question: when a large language reasoning model makes a choice, did it think first and then decide to, or decide first and then think? In this paper, we present evidence that detectable, early-encoded decisions shape chain-of-thought in reasoning models. Specifically, we show that a simple linear probe successfully decodes tool-calling decisions from pre-generation activations with very high confidence, and in some cases, even before a single reasoning token is produced. Activation steering supports this causally: perturbing the decision direction leads to inflated deliberation, and flips behavior in many examples (between 7 - 79% depending on model and benchmark). We also show through behavioral analysis that, when steering changes the decision, the chain-of-thought process often rationalizes the flip rather than resisting it. Together, these results suggest that reasoning models can encode action choices before they begin to deliberate in text.

大模型推理决策机制链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。