让AI裁判自己评自己,更准更可信。
A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth
- 用裁判的可靠度建模,避免偏见
- 比传统方法少用数据,结果更准
- 适合评估大模型排名与不确定性
在无标准答案的开放任务中,使用大语言模型作为裁判(LLM-as-a-judge)已成为主流。但不同裁判模型可靠性差异显著,若一视同仁会引发排行榜偏倚和误判。更多数据反而可能加剧错误。本文提出一种裁判感知的排序框架,基于改进的Bradley-Terry-Luce模型引入裁判特异性辨别参数,通过成对比较联合估计模型真实质量与裁判可靠性,无需参考标签。理论证明了可识别性、最大似然估计的一致性与渐近正态性,支持分数差与排名比较的置信区间。在多个公开基准与新构建数据集上,该方法提升与人类偏好的一致性,数据效率优于未加权基线,并实现校准的不确定性量化。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is that judge LLMs differ substantially in reliability; treating all judges equally can yield biased leaderboards and misleading uncertainty estimates. More data can make evaluation more confidently wrong under misspecified aggregation. We propose a judge-aware ranking framework that extends the Bradley-Terry-Luce model by introducing judge-specific discrimination parameters, jointly estimating latent model quality and judge reliability from pairwise comparisons without reference labels. We establish identifiability up to natural normalizations and prove consistency and asymptotic normality of the maximum likelihood estimator, enabling confidence intervals for score differences and rank comparisons. Across multiple public benchmarks and a newly collected dataset, our method improves agreement with human preferences, achieves higher data efficiency than unweighted baselines, and produces calibrated uncertainty quantification for LLM rankings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。