Chain-of-thought推理轨迹常在答案已定后继续生成,未必反映真实思考过程。
When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel

- 通过答案承诺代理检测每一步推理同步性,发现平均仅61.9%步骤一致。
- 58.0%的不一致源于答案已定仍持续生成推理文本,且答案未改变。
- 越依赖CoT的模型,推理轨迹越不忠实,但反而越有用,提示其为不可靠报告。
Chain-of-thought(CoT)轨迹被广泛用于提升语言模型能力与审计行为,隐含假设是可见轨迹始终与决定答案的计算过程同步。本文通过基于答案承诺代理的检测-分类-比较框架,在九个模型和七个推理基准上验证该假设。结果显示,潜藏承诺与显式答案到达仅在平均61.9%的步骤中对齐。主要不一致模式为“虚构延续”:58.0%的不一致事件发生在答案承诺已稳定之后,轨迹仍持续输出看似推理的内容,而空洞性分析表明此时答案未变化。在架构匹配的Qwen2.5与DeepSeek-R1-Distill对比中,推理管道变化比整体对齐度更显著影响失败模式,尤其在32B模型中,虚构步骤减少但矛盾状态增多。低步骤对齐度还与更高CoT效用相关,表明最受益于CoT的场景往往最缺乏时间上的忠实性。配对截断与捐赠者污染测试进一步表明,大量承诺后的文本并非答案形成的关键。这些发现说明,尽管CoT仍具实用性,但它作为答案形成时间的报告渠道并不可靠。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) traces are increasingly used both to improve language model capability and to audit model behavior, implicitly assuming that the visible trace remains synchronized with the computation that determines the answer. We test this assumption with a step-level Detect-Classify-Compare framework built around an answer-commitment proxy that is cross-validated with Patchscopes, tuned-lens probes, and causal direction ablation. Across nine models and seven reasoning benchmarks, latent commitment and explicit answer arrival align on only 61.9% of steps on average. The dominant mismatch pattern is confabulated continuation: 58.0% of detected mismatch events occur after the answer-commitment proxy has already stabilized while the trace continues producing deliberative-looking text, and a vacuousness analysis shows that the committed answer does not change during these steps. In architecture-matched Qwen2.5/DeepSeek-R1-Distill comparisons, the reasoning pipeline changes failure composition more than aggregate alignment, most clearly at 32B where confabulated steps decrease as contradictory states increase. Lower step-level alignment is also associated with larger CoT utility, suggesting that the settings that benefit most from CoT are often the least temporally faithful. Paired truncation and a complementary donor-corruption test further indicate that much post-commitment text is not load-bearing for the final answer. These findings suggest that CoT can remain useful while still being an unreliable report of when the answer was formed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。