LLM判官的偏见无法通过传统方法消除,需新设计才能可靠评估生成质量。
When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking

- 发现质量与偏见在成对比较中不可区分,数据无法恢复真实偏见
- 修正后得分依赖先验,而非数据,且仅在特定条件下有效
- 提出锚点门和配对渲染设计,适用于大规模低成本评估场景
大语言模型(LLMs)正被用作廉价、可扩展的判官,通过成对比较评估候选输出。由于这类判官偏好冗长或格式良好的回答,通常做法是引入偏见协变量并估计其影响。我们发现这种方法无效:质量与偏见的分离在成对比较中无法识别,且失败是精确的——在48个真实判官池中,系数的轮廓似然平坦至0.0000纳特;将比较规模扩大26倍也无改善。‘去偏’得分由先验决定,而非数据恢复。我们的贡献并非更好估计器,而是阐明何时基于先验的校正合理,并提供缺失信息时的解决方案。先验假设的质量与协变量无关,仅当相关性低于交叉点(0.22–0.60,取决于配置)时成立,这解释了同一模型在LLMBar上有效而在SummEval和Nectar上反而有害的原因。我们提出两种应对方案:可信锚点门(在6000次决策中无误触发,样本边界率≤6%),以及配对渲染设计。在十五个真实LLM判官中,偏见具有异质性且与能力相关:修正提升五位有偏但胜任的廉价判官的前K召回率0.20–0.32,对前沿判官无效(14位胜任判官中,能力与增益的斯皮尔曼ρ = -0.84,p < 10⁻³),使效益集中在规模化评估的关键位置。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise. Because such judges prefer verbose or well-formatted answers, the natural fix is to add bias covariates to a Bradley--Terry model and estimate the bias away. We show this cannot work as advertised: the quality/bias split is \emph{not identified} by pairwise comparisons, and the failure is exact -- across $48$ real judge-pools the profile likelihood over the coefficient is flat to $\mathbf{0.0000}$ \textbf{nats}, and scaling the comparisons $26\times$ buys none. A ``debiased'' score is selected by the prior, not recovered from data. Our contribution is accordingly not a better estimator but a characterization of \emph{when prior-based correction is justified}, plus designs that supply the missing information when it is not. The assumption the prior encodes -- quality is a priori uncorrelated with the covariate -- pays only while $\mathrm{corr}(θ,x)$ stays below a crossing point (configuration-dependent, $0.22$--$0.60$), which is what makes the same model help on LLMBar and hurt on SummEval and Nectar. We give two escapes: a \textbf{trusted-anchor gate} that decides per (judge, covariate, task) (no false enables in $6{,}000$ decisions at $K\ge10$ anchors, a rate our sample bounds at $\le6\%$), and a \textbf{paired rendering design}. Across fifteen real LLM judges bias is heterogeneous and capability-dependent: correction improves \topk{} recall by $0.20$--$0.32$ on five biased-but-competent cheap judges and is a no-op on frontier ones (Spearman $ρ{=}{-}0.84$ between competence and gain over the $14$ competent judges, $p{<}10^{-3}$), concentrating the benefit where at-scale evaluation happens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。