arXiv:2511.10201cs.CL2025-11被引 1

构建统一基准评估大模型高效推理,解决评测碎片化问题。

EffiReason-Bench: A Unified Benchmark for Evaluating and Advancing Efficient Reasoning in Large Language Models

  • 设计三类高效推理方法的统一评测框架
  • 在4个数据集上验证7种方法,发现无通用最优策略
  • 提出E3分数,兼顾效率与准确性的稳定评估

大语言模型使用思维链提示虽具强推理能力,但常产生冗长解释,增加成本且可能降低准确性。现有高效推理方法的公平比较受限于分散的评测实践。本文提出EffiReason-Bench,一个统一基准,用于跨范式评估三类高效推理方法:推理蓝图、动态执行与事后精炼。为实现逐步评估,通过标准化结构、全面选项分析及人工验证,构建了CommonsenseQA与LogiQA的经验证思维链标注。我们在4个涵盖数学、常识与逻辑的数据集上,对6个开源大模型(1B-70B参数)上的7种方法进行了评估,并提出受经济权衡建模启发的E3-Score,该指标平滑稳定,无需依赖启发式规则。实验表明,无单一方法在所有场景下占优;最优策略取决于模型规模、任务复杂度与架构。

原文摘要 · Abstract (English)

Large language models (LLMs) with Chain-of-Thought (CoT) prompting achieve strong reasoning but often produce unnecessarily long explanations, increasing cost and sometimes reducing accuracy. Fair comparison of efficiency-oriented approaches is hindered by fragmented evaluation practices. We introduce EffiReason-Bench, a unified benchmark for rigorous cross-paradigm evaluation of efficient reasoning methods across three categories: Reasoning Blueprints, Dynamic Execution, and Post-hoc Refinement. To enable step-by-step evaluation, we construct verified CoT annotations for CommonsenseQA and LogiQA via a pipeline that enforces standardized reasoning structures, comprehensive option-wise analysis, and human verification. We evaluate 7 methods across 6 open-source LLMs (1B-70B) on 4 datasets spanning mathematics, commonsense, and logic, and propose the E3-Score, a principled metric inspired by economic trade-off modeling that provides smooth, stable evaluation without discontinuities or heavy reliance on heuristics. Experiments show that no single method universally dominates; optimal strategies depend on backbone scale, task complexity, and architecture.

大模型推理评测基准思维链效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。