大模型当裁判评估存在偏差,新方法可诊断可靠性问题。
Bias and Uncertainty in LLM-as-a-Judge Estimation
- 用法官质量与跨模型校准不稳定性作为诊断指标
- 共享校准可能导致判断方向错误且看似可信
- 适合关注模型评估可靠性的研究者参考
大模型作为裁判(LLM-as-a-Judge)已成为评估基础模型性能的标准工具。然而,直接使用原始裁判输出的朴素估计量会系统性地产生偏差。尽管已有研究提出校正偏差的估计方法,但其可靠性高度依赖于裁判质量(J)以及模型比较中的校准稳定性(ΔJ)。实际中共享校准虽具吸引力,却可能引入严重偏差,甚至在某些情况下导致比较结果方向相反且具有高置信度。本文通过理论分析、针对裁判质量(J)和跨模型校准不稳定性(ΔJ)的模拟,以及基于真实数据的MMLU-Pro案例研究,揭示了此类失效模式。我们提出将J和ΔJ作为校正估计量(尤其是共享校准比较)不可靠时的诊断依据,并为LaaJ评估提供报告规范。
原文摘要 · Abstract (English)
LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed estimators to correct this bias, but their reliability depends critically on judge quality and, for model comparisons, on calibration stability. Sharing calibration across compared models is practically attractive but can introduce severe bias, including cases where the comparison estimate points in the wrong direction with high apparent confidence. We study these failure modes through analytical results, simulations over judge quality ($J$) and cross-model calibration instability ($ΔJ$), and a real-data MMLU-Pro case study with sign reversal. We propose $J$ and $ΔJ$ as diagnostics for when corrected estimates, especially shared-calibration comparisons, are likely unreliable, and provide reporting guidance for LaaJ evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。