用轮换评委法消除大模型评估中的系统性偏差
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
- 采用轮换评委机制,避免单个评委偏好影响结果
- 在MT-Bench和MindEval上验证,评分准确率显著提升
- 适合需要公平评估的通用与垂直领域研究者
基于大模型作为评判者的评估已成为开放生成任务的标准方法;然而,评委存在系统性偏差,无法通过增加场景或生成次数来消除。这些偏差的幅度常与模型差异相当,导致单评委评估结果不可靠。本文提出一种方差分解方法,将基准得分方差拆分为场景、生成、评委和残差成分。基于此分析,提出循环评委分配策略(CyclicJudge),在固定评委数量和评估成本下,可精确恢复评委组平均分,且计算开销与单评委评估相同。在MT-Bench和MindEval上的实证结果验证了该方法的有效性,适用于通用及特定领域评估场景。
原文摘要 · Abstract (English)
LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are often similar in magnitude to the model differences that benchmarks are designed to detect, resulting in unreliable rankings when single-judge evaluations are used. We introduce a variance decomposition that partitions benchmark score variance into scenario, generation, judge, and residual components. Based on this analysis, CyclicJudge, a round-robin assignment of judges to scenarios, is demonstrated to be the optimal strategy for a fixed judge panel and judge-call budget: the score recovers the panel mean exactly while matching the cost of single-judge evaluation. Empirical results on MT-Bench and MindEval validate the effectiveness of CyclicJudge as predicted, across both general-purpose and domain-specific evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。