arXiv:2510.09738cs.CLcs.AI2025-10被引 25

用新方法评估大模型当评分员的能力,发现大小不是关键,训练策略才决定评分是否像人。

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

  • 分两步评估:先看相关性,再用z值区分人类式和过度一致的评分模式。
  • 54个模型中23个像人一样有自然波动,4个过于一致可能太死板或太可靠。
  • 不靠模型大小,而是训练策略决定大模型能否像人一样判断回答好坏。

本研究提出Judge's Verdict基准,采用两步法评估54个大语言模型作为评分员在RAG或智能体流水线响应准确性评估中的表现。通过从传统相关性分析转向全面的Cohen's Kappa分析,该方法识别出两种判断模式:人类式判断(|z| < 1),体现自然的人类差异;超一致性判断(z > 1),超越典型人与人之间的同意水平。结果表明,27个模型达到一级性能:23个模型表现出人类式判断,保留了人类评判的细微差别;4个模型呈现超一致性行为,可能反映更高可靠性或对复杂判断的过度简化。测试涵盖43个开源模型(参数量1B-405B)和11个闭源模型(GPT、Gemini、Claude系列),揭示评分能力并非仅取决于模型规模,而与特定训练策略密切相关。主要贡献包括:(1)证明相关性不足以评估评分者;(2)提出基于判断模式的“法官图灵测试”;(3)提供标准化基准,可将LLM评分员按不同需求分类至不同性能层级。

原文摘要 · Abstract (English)

This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We assess how well 54 LLMs can replicate human judgment when scoring responses from RAG (Retrieval-Augmented Generation) or Agentic pipelines against ground truth answers. Our methodology progresses from traditional correlation analysis to comprehensive Cohen's Kappa analysis that measures actual agreement patterns. The two-step approach includes: (1) a correlation test that filters judges with strong alignment, followed by (2) a human-likeness test using z-scores to identify two distinct judgment patterns: human-like judgment (|z| < 1) that mimics natural human variation, and super-consistent judgment (z > 1) that exceeds typical human-to-human agreement levels. This methodology reveals that 27 out of 54 tested LLMs achieve Tier 1 performance: 23 models exhibit human-like patterns that preserve the nuances of human judgment, while 4 models demonstrate super-consistent behavior, a pattern that could indicate either enhanced reliability or oversimplification of complex judgments. Testing 43 open-source models (1B-405B parameters) and 11 closed models (GPT, Gemini, Claude variants), we demonstrate that judge excellence is not solely dependent on model size but on specific training strategies. Our key contributions include: (1) establishing that correlation alone is insufficient for judge evaluation, (2) introducing a "Turing Test for judges" based on agreement patterns, and (3) providing a standardized benchmark for classifying LLM judges into distinct performance tiers for different evaluation needs.

大模型评测判断一致性人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。