arXiv:2601.03550cs.AI2026-01ACL被引 2

提出新评估框架,揭示大模型推理效率的真正瓶颈

ReEfBench: Quantifying the Reasoning Efficiency of LLMs

  • 构建神经符号框架,无侵入式分析推理过程
  • 发现四种行为原型,证明长推理非深度思考必要条件
  • 适合研究模型推理机制与训练策略的学者

测试时扩展使大语言模型能够处理复杂推理,但现有思维链(CoT)评估方法难以区分性能提升是源于真实推理还是单纯冗长。为此,(1) 我们提出一种新型神经符号框架,实现无侵入、全面的过程中心型推理评估;(2) 借此识别出四种不同行为原型并诊断失败模式;(3) 分析了推理模式、训练策略与模型规模的影响。结果表明,深度推理并不依赖于长文本生成。此外,我们发现:训练中混合长短思维链数据会导致过早饱和与崩溃;将模型蒸馏至小模型虽保留行为长度,但因容量限制无法复制逻辑有效性。

原文摘要 · Abstract (English)

Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performance gains stem from genuine reasoning or mere verbosity. To address this, (1) we propose a novel neuro-symbolic framework for the non-intrusive, comprehensive process-centric evaluation of reasoning. (2) Through this lens, we identify four distinct behavioral prototypes and diagnose the failure modes. (3) We examine the impact of inference mode, training strategy, and model scale. Our analysis reveals that extended token generation is not a prerequisite for deep reasoning. Furthermore, we reveal critical constraints: mixing long and short CoT data in training risks in premature saturation and collapse, while distillation into smaller models captures behavioral length but fails to replicate logical efficacy due to intrinsic capacity limits.

推理评估大模型思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。