arXiv:2502.12289cs.CL2025-02EMNLP综述被引 66

系统梳理大模型推理过程评估方法,提出四维评价体系。

Evaluating Step-by-step Reasoning Traces: A Survey

  • 构建事实性、有效性、连贯性、实用性四维评估框架
  • 分析多类数据集与评测工具,揭示当前评估不一致问题
  • 为提升大模型推理能力研究提供可参考的评估方向

逐步推理被广泛用于提升大语言模型在复杂任务中的推理能力。评估推理轨迹的质量对于理解与改进大模型的推理至关重要。然而,现有的评估方法极不统一,导致评估器设计与基准测试发展呈现碎片化。为填补这一空白,本文对逐步推理评估进行了全面综述,提出了包含事实性、有效性、连贯性和实用性四个顶层类别的评估标准分类体系。基于该分类,我们回顾了不同数据集、评估器实现及最新研究发现,指明未来研究的有前景方向。

原文摘要 · Abstract (English)

Step-by-step reasoning is widely used to enhance the reasoning ability of large language models (LLMs) in complex problems. Evaluating the quality of reasoning traces is crucial for understanding and improving LLM reasoning. However, existing evaluation practices are highly inconsistent, resulting in fragmented progress across evaluator design and benchmark development. To address this gap, this survey provides a comprehensive overview of step-by-step reasoning evaluation, proposing a taxonomy of evaluation criteria with four top-level categories (factuality, validity, coherence, and utility). Based on the taxonomy, we review different datasets, evaluator implementations, and recent findings, leading to promising directions for future research.

推理评估大模型评测体系LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。