arXiv:2604.23178cs.AI2026-04被引 5

对比九种去偏策略,发现低成本模型加去偏可超顶级大模型评估效果。

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

论文配图:Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
图 1 · 摘自论文原文
  • 用九种策略在五种模型上系统测试去偏效果
  • 中端模型+组合预算策略达成71.0%一致率,成本仅顶级模型的1/15
  • 首次揭示风格偏见主导评估,且不同模型对冗长响应反应各异

LLM作为裁判已成为语言模型输出评估的主流范式,但其存在系统性偏差,影响评估可靠性。本文开展全面实证研究,对比九种去偏策略在四种厂商(Google、Anthropic、OpenAI、Meta)提供的五种模型上的表现,涵盖三个基准(MT-Bench n=400,LLMBar n=200,自建数据集n=375)和四类偏差。核心发现:采用合适去偏策略的中端模型(Gemini 2.5 Flash + 组合预算策略)达到最高一致性(71.0%,kappa=0.549),每评估成本约$0.001,仅为最佳前沿配置(Claude Sonnet 4,69.5%)的1/15。风格偏见最显著(0.10–0.76),远超位置偏见(≤0.04),但极少被研究;在长度感知下,各模型对冗长回应反应不一:Pro/Flash/Llama偏好更长答案(+0.24至+0.44),Claude偏好简洁(-0.12),GPT-4o中立(-0.04);在截断控制下,所有模型均正确偏好完整回答(准确率0.88–1.00)。去偏策略提升多模型性能:Claude S8(+11.5 pp)、Flash S8(+7.5 pp)、Claude S5(+7.3 pp)通过霍尔姆-邦费罗尼校正,其余亦显著。研究发布评估框架、375对受控数据集及所有九种策略的缓存结果。

原文摘要 · Abstract (English)

LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM judges exhibit systematic biases that compromise evaluation reliability. We present a comprehensive empirical study comparing nine debiasing strategies across five judge models from four provider families (Google, Anthropic, OpenAI, Meta), three benchmarks (MT-Bench n=400, LLMBar n=200, custom n=375), and four bias types. Our headline practical finding is that a mid-tier model with the right debiasing can outperform frontier judges at a fraction of the cost: Gemini 2.5 Flash with the Combined Budget strategy reaches the highest agreement of any configuration we tested (71.0%, kappa=0.549) at ~$0.001 per evaluation, about 15x cheaper than the best frontier setup (Claude Sonnet 4, 69.5%, ~$0.015). Other key findings: (1) Style bias is the dominant bias (0.10-0.76 across models, favoring markdown over plain prose), far exceeding position bias (<=0.04), yet is rarely studied. (2) Verbosity bias is heterogeneous when measured length-aware: Pro, Flash, and Llama prefer longer answers (+0.24 to +0.44), Claude prefers concise (-0.12), and GPT-4o is neutral (-0.04); on truncation controls all models correctly prefer the complete response (0.88-1.00 accuracy). (3) Debiasing helps multiple models: Claude S8 (+11.5 pp), Flash S8 (+7.5 pp), and Claude S5 (+7.3 pp) survive Holm-Bonferroni correction, with Flash S1 (+4.7 pp) and Llama S8 (+4.5 pp) also significant. We release our evaluation framework, the 375-pair controlled dataset, and per-instance cached results for all nine strategies.

大模型评估去偏策略评估可靠性成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。