arXiv:2605.09702stat.MEcs.CL2026-05被引 3

用全部模型评判者比挑最准的更准,关键在校准而非筛选。

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

  • 保留所有评判者,通过校准提升评估准确性
  • 全面板在RewardBench2上NLL低至0.006,是选前5的半数
  • 即使表现低于随机,只要偏差可学且不冗余就仍有价值

多评判者评估被广泛用于大模型与奖励模型的评测,传统做法是筛选:只保留最准确的评判者,剔除较弱者。本文指出,当目标是基于标注校准集进行概率校准评估时,这一策略可能适得其反。固定聚合与校准方法,对比按准确率选取前k名评判者与使用全部评判者的效果。在四个涵盖大模型作为评判者和奖励模型设置的成对评估基准上,全面板始终优于按准确率筛选。在RewardBench2上,全面板实现0.006的负对数似然(NLL),而前5名筛选仅得0.013,校准误差减半。该优势在去重同类评判者家族及更强子集搜索下依然成立。通过理想分析揭示:在合理评分规则下,额外评判信号不会增加最优校准风险;即使表现低于随机,只要偏差可学习、信号非冗余,仍具价值。核心原则为:有校准数据时,多评判者评估不应仅按准确度淘汰弱者,应保留可解析、非冗余、可校准的评判者。

原文摘要 · Abstract (English)

Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges and discard weaker ones. We show that this heuristic can reverse when the target is not point accuracy, but calibrated probabilistic evaluation from a labeled calibration set. Holding the aggregation and calibration procedures fixed, we compare accuracy-ranked top-$k$ judge selection with using the full judge panel. Across four labeled pairwise-evaluation benchmarks spanning LLM-as-judge and reward-model settings, the calibrated full panel consistently outperforms accuracy-based selection. On RewardBench2, retaining all judges achieves negative log-likelihood (NLL) of $0.006$ versus $0.013$ under top-5 selection, halving the calibration error. This advantage persists after judge-family deduplication and against stronger same-pipeline subset search. We explain this reversal with oracle analyses showing that the optimal calibrated risk under proper scoring rules cannot increase when additional judge signals are made available, and that even below-chance judges can be useful when their biases are learnable and their signals are non-redundant. The resulting operating principle is simple: in multi-judge evaluation with labeled calibration data, do not discard weak judges by accuracy alone; keep them when they are parseable, non-redundant, and calibratable.

大模型评测校准评估多评判者

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。