评测大模型裁判的可靠性,发现强模型表现仅略优于随机
JudgeBench: A Benchmark for Evaluating LLM-based Judges
- 构建新评估框架,用客观正确性标签生成挑战性对比样本
- 多模型测试显示GPT-4o等仍仅略高于随机猜测水平
- 适合评估高阶大模型裁判,尤其在知识推理任务中
基于大模型的裁判已成为替代人工评估的可扩展方案,广泛用于模型评估与优化。然而,这类裁判自身的可靠性却极少被检验。随着大模型能力提升,其输出愈发复杂,需更强裁判进行评估。现有基准主要关注与人类偏好对齐,但难以应对人类众包偏好无法反映事实与逻辑正确性的难题。为此,本文提出全新评估框架,并构建JudgeBench基准,涵盖知识、推理、数学与编码四类挑战性响应对。该基准通过新流水线将现有困难数据集转换为带有客观正确性偏好标签的对比样本。对提示型裁判、微调裁判、多智能体裁判及奖励模型的全面评估表明,相比以往基准,JudgeBench显著提升难度,许多强模型(如GPT-4o)表现仅略优于随机猜测。整体而言,JudgeBench为评估日益先进的大模型裁判提供了可靠平台。数据与代码已公开于https://github.com/ScalerLab/JudgeBench。
原文摘要 · Abstract (English)
LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requiring stronger judges to evaluate them. Existing benchmarks primarily focus on a judge's alignment with human preferences, but often fail to account for more challenging tasks where crowdsourced human preference is a poor indicator of factual and logical correctness. To address this, we propose a novel evaluation framework to objectively evaluate LLM-based judges. Based on this framework, we propose JudgeBench, a benchmark for evaluating LLM-based judges on challenging response pairs spanning knowledge, reasoning, math, and coding. JudgeBench leverages a novel pipeline for converting existing difficult datasets into challenging response pairs with preference labels reflecting objective correctness. Our comprehensive evaluation on a collection of prompted judges, fine-tuned judges, multi-agent judges, and reward models shows that JudgeBench poses a significantly greater challenge than previous benchmarks, with many strong models (e.g., GPT-4o) performing just slightly better than random guessing. Overall, JudgeBench offers a reliable platform for assessing increasingly advanced LLM-based judges. Data and code are available at https://github.com/ScalerLab/JudgeBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。