arXiv:2604.18307cs.CL2026-04被引 1

模型内部激活值能捕捉推理步骤的重要性,比文本本身更可靠。

Reasoning Models Know What's Important, and Encode It in Their Activations

  • 用激活值训练探测器识别推理步骤重要性
  • 不同模型对重要步骤判断高度一致
  • 适合研究模型内部推理机制的学者

语言模型解决复杂任务时常生成多步推理链,其中部分步骤至关重要,其余可删减。确定哪些步骤真正关键仍是一个核心问题。我们考察应从模型内部还是推理链文本本身入手。结果表明,模型激活值包含的信息量远超文本 tokens,能更准确识别关键步骤。通过在激活值上训练探测器,我们发现模型在生成后续步骤前已编码了步骤重要性的内部表示。不同模型的内部表示对重要步骤的判断高度一致,且分布于多层,与位置或长度等表面特征无关。这说明仅看文本会遗漏深层推理机制,推理分析必须深入模型内部。

原文摘要 · Abstract (English)

Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the final answer, others are removable. Determining which steps matter most, and why, remains an open question central to understanding how models process reasoning. We investigate if this question is best approached through model internals or through tokens of the reasoning chain itself. We find that model activations contain more information than tokens for identifying important reasoning steps. Crucially, by training probes on model activations to predict importance, we show that models encode an internal representation of step importance, even prior to the generation of subsequent steps. The internal representations of importance in different models yield high agreement on which steps are important. The representation is distributed across layers, and does not correlate with surface-level features, such as a step's relative position or its length. Our findings suggest that analyzing activations can reveal aspects of reasoning that surface-level approaches fundamentally miss, indicating that reasoning analyses should look into model internals.

模型推理激活分析重要性识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。