arXiv:2601.00514cs.AIcs.CL2026-01ACL被引 10

模型的突然顿悟其实只是推理不稳的表现,而非真正智能。

The Illusion of Insight in Reasoning Models

  • 分析上百万推理轨迹,发现模型中途改变思路极少见
  • 这些转变很少提升准确率,且随训练不变得更频繁
  • 高不确定性时人为触发改变反而能提高精度

推理模型是否具备‘顿悟’时刻?以往研究认为像DeepSeek-R1-Zero这样的模型会在推理过程中突然调整策略并产出正确结果,暗示其具有自我修正能力。然而,这种内在策略转变是否真能提升性能仍不明确。本文分析了超过100万条推理轨迹、数百个训练检查点、三个推理领域及多种解码温度与模型架构。结果显示,中段推理转变极为罕见,不会随训练增加,且极少提升准确率,表明其并非真正意义上的模型洞察。但其影响受模型不确定性的调节。基于此,我们发现,在高熵状态下人为触发外在策略转变可稳定提升准确率。结论表明,中段推理转变是不稳定的推理表现,而非自我修正的内在机制。

原文摘要 · Abstract (English)

Do reasoning models have "Aha!" moments? Prior work suggests that models like DeepSeek-R1-Zero undergo sudden mid-trace realizations that lead to accurate outputs, implying an intrinsic capacity for self-correction. Yet, it remains unclear whether such intrinsic shifts in reasoning strategy actually improve performance. Here, we study mid-reasoning shifts and instrument training runs to detect them. Our analysis spans 1M+ reasoning traces, hundreds of training checkpoints, three reasoning domains, and multiple decoding temperatures and model architectures. We find that reasoning shifts are rare, do not become more frequent with training, and seldom improve accuracy, indicating that they do not correspond to prior perceptions of model insight. However, their effect varies with model uncertainty. Building on this finding, we show that artificially triggering extrinsic shifts under high entropy reliably improves accuracy. Our results show that mid-reasoning shifts are symptoms of unstable inference behavior rather than an intrinsic mechanism for self-correction.

推理模型模型洞察稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。