arXiv:2604.10739cs.AI2026-04ACL被引 10

发现大模型推理越想越错,适度思考更高效

When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling

  • 测试时扩展思维链,但过长反而降低准确率
  • 高算力下边际收益锐减,出现'过度思考'现象
  • 不同难度题应配不同思考长度,适合动态调度

通过延长思维链提升大语言模型推理能力已成为主流方法。然而,现有研究隐含假设:更长的思考总能带来更好结果,这一假设尚未被验证。本文系统研究了随着计算预算增加,额外推理令牌的边际效用变化。结果表明,高预算下边际收益显著下降,模型出现'过度思考',即延长推理导致放弃原本正确的答案。此外,最优思考长度随题目难度变化,说明统一分配计算资源并非最优。成本敏感评估框架显示,在中等预算停止推理可大幅减少计算量,同时保持相近准确率。

原文摘要 · Abstract (English)

Scaling test-time compute through extended chains of thought has become a dominant paradigm for improving large language model reasoning. However, existing research implicitly assumes that longer thinking always yields better results. This assumption remains largely unexamined. We systematically investigate how the marginal utility of additional reasoning tokens changes as compute budgets increase. We find that marginal returns diminish substantially at higher budgets and that models exhibit ``overthinking'', where extended reasoning is associated with abandoning previously correct answers. Furthermore, we show that optimal thinking length varies across problem difficulty, suggesting that uniform compute allocation is suboptimal. Our cost-aware evaluation framework reveals that stopping at moderate budgets can reduce computation significantly while maintaining comparable accuracy.

大模型推理思维链效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。