arXiv:2409.04168cs.CLcs.AI2024-09被引 25

用LLM当数学题裁判,发现它偏爱好模型但会误判。

From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks

  • 通过多步推理任务测试LLM裁判表现,发现难样本难判断。
  • 裁判表现与模型质量强相关,常误判错误答案为正确。
  • 仅用词性标签就可预测70%-75%的裁判决策,适合自动化评估。

为减少人工标注依赖,大语言模型(LLMs)被提出作为其他候选模型生成质量的评判者。现有评估通常基于摘要或机器翻译任务中与人类判断的相关性。本文聚焦于数学推理任务上的LLM裁判表现。此类任务需多步推理且答案可验证,便于客观评估。我们进行详细分析发现:简单样本易判断,复杂样本则困难。分析揭示裁判表现与候选模型任务性能高度相关,表明裁判倾向于偏好高质量模型,即使其答案错误。为此,我们测试是否可用词性标签等简单特征预测裁判行为,结果表明可正确预测70%–75%的判断。最后,通过实际应用案例分析显示,LLM裁判能稳定识别平均更优的模型,但在提升任务性能方面效果有限。

原文摘要 · Abstract (English)

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human judgments on generative tasks such as summarization or machine translation. In contrast, we study LLM judges on mathematical reasoning tasks. These tasks require multi-step reasoning, and the correctness of their solutions is verifiable, enabling a more objective evaluation. We perform a detailed performance analysis and find that easy samples are easy to judge, and difficult samples are difficult to judge. Our analysis uncovers a strong correlation between judgment performance and the candidate model task performance, indicating that judges tend to favor higher-quality models even if their answer is incorrect. As a consequence, we test whether we can predict the behavior of LLM judges using simple features such as part-of-speech tags and find that we can correctly predict 70%-75% of judgments. We conclude this study by analyzing practical use cases, showing that LLM judges consistently detect the on-average better model but largely fail if we use them to improve task performance.

LLM裁判数学推理自动评估模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。