提出新方法检测推理链是否真实影响模型决策
Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning
- 通过破坏推理步骤观察置信度变化,量化每步重要性
- 发现推理有效范围仅占链长70%-85%,之后步骤基本无效
- 揭示模型可有正确内部表示却仍失败,适合可信AI研究者
链式思维(CoT)解释被广泛用于说明语言模型如何解决复杂问题,但其是否真实反映模型决策过程尚不明确。本文提出归一化对数差异衰减(NLDD)指标,通过扰动推理链中各步骤并测量模型对答案置信度的下降程度,判断该步骤是否真正关键。通过标准化处理,NLDD实现跨不同架构模型的严谨比较。在句法、逻辑和算术任务中测试三种模型家族,发现推理有效边界(k*)稳定在链长的70%至85%之间,超出后推理词元对最终答案影响微弱甚至为负。此外,模型可在内部保持正确表征却完全失败任务。结果表明,仅凭准确率无法判断模型是否真正进行推理。NLDD为评估链式思维实际作用提供了工具。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) explanations are widely used to interpret how language models solve complex problems, yet it remains unclear whether these step-by-step explanations reflect how the model actually reaches its answer, or merely post-hoc justifications. We propose Normalized Logit Difference Decay (NLDD), a metric that measures whether individual reasoning steps are faithful to the model's decision-making process. Our approach corrupts individual reasoning steps from the explanation and measures how much the model's confidence in its answer drops, to determine if a step is truly important. By standardizing these measurements, NLDD enables rigorous cross-model comparison across different architectures. Testing three model families across syntactic, logical, and arithmetic tasks, we discover a consistent Reasoning Horizon (k*) at 70--85% of chain length, beyond which reasoning tokens have little or negative effect on the final answer. We also find that models can encode correct internal representations while completely failing the task. These results show that accuracy alone does not reveal whether a model actually reasons through its chain. NLDD offers a way to measure when CoT matters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。