arXiv:2501.14492cs.CLcs.AI2025-01被引 11

新基准评估大模型批评能力,发现高级模型显著优于传统模型。

RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques

  • 采用闭环评估法,通过修正效果衡量批评质量。
  • o1-mini在所有批评场景中均显著优于经典模型。
  • 自批评与迭代批评中,传统模型甚至表现更差。

批评对提升大语言模型性能至关重要,可实现自我改进和他人反馈。然而,由于任务开放性,评估模型的批评能力极具挑战。本文提出新基准,通过闭环方法评估批评生成的修正质量。与现有开环基准不同,该基准包含自批评、交叉批评和迭代批评,能有效区分先进推理模型与传统模型的能力。基于八个复杂推理任务进行实验。结果表明:尽管在直接链式思维生成上表现相近,经典模型在所有批评场景中均显著落后于先进模型o1-mini;在自批评与迭代批评设置下,传统模型甚至低于基线表现。代码与数据已公开。

原文摘要 · Abstract (English)

Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and suggesting improvements. However, evaluating the critique capabilities of LLMs presents a significant challenge due to the open-ended nature of the task. In this work, we introduce a new benchmark designed to assess the critique capabilities of LLMs. Unlike existing benchmarks, which typically function in an open-loop fashion, our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques. Moreover, the benchmark incorporates features such as self-critique, cross-critique, and iterative critique, which are crucial for distinguishing the abilities of advanced reasoning models from more classical ones. We implement this benchmark using eight challenging reasoning tasks. We have several interesting findings. First, despite demonstrating comparable performance in direct chain-of-thought generation, classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios. Second, in self-critique and iterative critique settings, classical LLMs may even underperform relative to their baseline capabilities. We hope that this benchmark will serve as a valuable resource to guide future advancements. The code and data are available at \url{https://github.com/tangzhy/RealCritic}.

大模型评估批评机制闭环评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。