评测大模型推理效率,发现多数模型在简单任务中过度思考
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
- 构建新基准Think-Bench,量化评估模型推理过程的效率
- 发现多数模型在简单问题上生成冗长无用的推理链
- 适合关注推理效率与推理质量平衡的研究者
大型推理模型(LRMs)在复杂任务中表现优异,常超越传统大语言模型(LLMs)。然而,普遍存在的过度思考问题严重制约其计算效率。过度思考指模型生成大量冗余且对准确结果贡献极小的文本,尤其在简单任务中,造成显著算力浪费。为此,我们提出Think-Bench基准,用于系统评估LRMs的推理效率。我们设计了新的效率指标,并对多种LRMs进行了多维度综合评估,涵盖推理过程、结果质量及思维链(CoT)特性。分析显示,多数LRMs在处理简单问题时存在过度思考,生成不必要的长推理链。尽管许多模型具备高质量的CoT,但效率普遍偏低。我们希望Think-Bench能为推进LRMs研究提供坚实基础。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) have achieved impressive performance in complex tasks, often outperforming conventional large language models (LLMs). However, the prevalent issue of overthinking severely limits their computational efficiency. Overthinking occurs when models generate excessive and redundant tokens that contribute little to accurate outcomes, especially in simple tasks, resulting in a significant waste of computational resources. To systematically investigate this issue, we introduce Think-Bench, a benchmark designed to evaluate the reasoning efficiency of LRMs. We also propose novel efficiency metrics and conduct a comprehensive evaluation of various LRMs across multiple dimensions, including the reasoning process, outcome quality, and chain-of-thought (CoT) characteristics. Our analysis reveals that most LRMs exhibit overthinking in handling easy questions, generating unnecessarily lengthy reasoning chains. While many LRMs demonstrate high CoT quality, several suffer from low efficiency. We hope that Think-Bench can serve as a robust foundation for advancing research into LRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。