arXiv:2508.13141cs.CLcs.LG2025-08被引 38

评测大模型在思考过与不足间的平衡,发现无一模型能兼顾效率与准确。

OptimalThinkingBench: Evaluating Over and Underthinking in LLMs

  • 构建双子基准:一个测过度思考,一个测思考不足。
  • 33个模型测试显示,大模型常为简单问题多算数百词却无效。
  • 小模型虽快但难解难题,提示需更智能的思考决策机制。

思维型大模型在复杂任务中表现优异,但消耗更多计算资源,对简单问题过度思考;非思维型模型虽快速廉价,却在高难度推理中思考不足。这促使研究者发展了独立的思维与非思维模型,使用户需自行选择最优模型。我们提出统一基准OptimalThinkingBench,联合评估过度思考与思考不足问题,并推动开发性能与效率平衡的最优思考模型。该基准包含两个子基准:OverthinkingBench(72个领域内的简单数学与通用查询)和UnderthinkingBench(11个高难度推理任务及更复杂的数学题)。通过新型思考调整精度指标,我们评估了33种不同思维与非思维模型,发现无一模型能在本基准上实现最优思考。思维模型常在最简单查询上进行数百词的冗余计算,性能未提升;而大型非思维模型则因思考不足,表现远逊于小型思维模型。我们进一步探索多种促进最优思考的方法,但发现这些方法往往在某一子基准上提升,却损害另一子基准表现,凸显未来亟需更统一、更智能的思考模型。

原文摘要 · Abstract (English)

Thinking LLMs solve complex tasks at the expense of increased compute and overthinking on simpler problems, while non-thinking LLMs are faster and cheaper but underthink on harder reasoning problems. This has led to the development of separate thinking and non-thinking LLM variants, leaving the onus of selecting the optimal model for each query on the end user. We introduce OptimalThinkingBench, a unified benchmark that jointly evaluates overthinking and underthinking in LLMs and also encourages the development of optimally-thinking models that balance performance and efficiency. Our benchmark comprises two sub-benchmarks: OverthinkingBench, featuring simple math and general queries in 72 domains, and UnderthinkingBench, containing 11 challenging reasoning tasks along with harder math problems. Using novel thinking-adjusted accuracy metrics, we extensively evaluate 33 different thinking and non-thinking models and show that no model is able to optimally think on our benchmark. Thinking models often overthink for hundreds of tokens on the simplest user queries without improving performance. In contrast, large non-thinking models underthink, often falling short of much smaller thinking models. We further explore several methods to encourage optimal thinking, but find that these approaches often improve on one sub-benchmark at the expense of the other, highlighting the need for better unified and optimal models in the future.

大模型评估推理能力效率平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。