arXiv:2512.16041cs.CLcs.AI2025-12被引 20

提出无须人工标注的LLM裁判评估框架,发现顶级模型在判断时仍不一致。

Are We on the Right Way to Assessing LLM-as-a-Judge?

  • 用理性选择理论设计双维度评估:局部自洽性与全局逻辑一致性
  • 在650个真实问题上测试,顶级模型在近四分之一难题中偏好不稳
  • 揭示人类判断也存在显著不一致,挑战人工标注的金标准地位

LLM-as-a-Judge 被广泛用于模型评估和训练中的监督奖励。然而现有基准主要依赖人工标注,引入偏见并限制可扩展性。为此,我们提出 Sage,一个无需任何人工标注即可评估 LLM 裁判质量的新评测套件。受理性选择理论启发,Sage 引入两个新视角:局部自洽性(成对偏好稳定性)和全局逻辑一致性(完整偏好集的传递性)。我们通过结合结构化基准题与真实用户查询,构建了包含650个问题的数据集。实验表明,该指标具有高稳定性且与 LLMBar、RewardBench2 等监督基准高度相关,验证了 Sage 的可靠性。基于 Sage 发现,当前最先进模型在评分与成对比较任务中均存在显著可靠性问题;即使顶级模型 Gemini-2.5-Pro 与 GPT-5,在约四分之一困难案例中也无法保持一致偏好。我们将其归因于一种新现象——情境偏好,说明明确的评分标准有助于模型保持一致性。进一步分析表明,微调 LLM-as-a-Judge 可提升性能,采用多人评审或深度推理也能增强判断一致性。此外,人类判断也存在大量不一致,暗示人工标注未必是可靠金标准。

原文摘要 · Abstract (English)

LLM-as-a-Judge has been widely adopted as an evaluation method and served as supervised rewards in model training. However, existing benchmarks for LLM-as-a-Judge are mainly relying on human-annotated ground truth, which introduces human bias that undermines the assessment of reliability and imposes scalability constraints. To overcome these limitations, we introduce Sage, a novel evaluation suite that assesses the quality of LLM judges without necessitating any human annotation. Inspired by axioms of rational choice theory, Sage introduces two new lenses for measuring LLM-as-a-Judge: local self-consistency (pair-wise preference stability) and global logical consistency (transitivity across a full set of preferences). We curate a dataset of 650 questions by combining structured benchmark problems with real-world user queries. Our experiments demonstrate both the stability of our metrics and their high correlation with supervised benchmarks like LLMBar and RewardBench2, confirming Sage's reliability as an evaluation suite for the robustness and accuracy of LLM-as-a-Judge. Based on Sage, we reveal that current state-of-the-art LLMs exhibit significant reliability problems when acting as judges in both scoring and pairwise settings; even the top-performing models, Gemini-2.5-Pro and GPT-5, fail to maintain consistent preferences in nearly a quarter of difficult cases. We attribute this to a new phenomenon called situational preference, which explains why explicit rubrics or criteria can help the model judge consistently across answer pairs. Our further analysis shows that finetuned LLM-as-a-Judge is a feasible method to boost performance, and the panel-based judge as well as deep reasoning can enhance the judging consistency. We also find substantial inconsistency in human judgments, which indicates that human annotation may not be a reliable gold standard.

大模型评估裁判模型一致性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。