arXiv:2604.22891cs.LGcs.AI2026-04被引 8

发现大模型评自分输出有偏见,提出自动纠正方法

Quantifying and Mitigating Self-Preference Bias of LLM Judges

论文配图:Quantifying and Mitigating Self-Preference Bias of LLM Judges
图 1 · 摘自论文原文
  • 构建质量相近的输出对,自动分离评价偏差与判断力
  • 实测20个主流大模型,能力越强反而越易自我偏袒
  • 设计分步评估策略,平均降低31.5%评价偏差

LLM作为评价者已成为自动化评估系统的核心,广泛应用于模型对齐、排行榜构建和质量控制等场景。然而,自偏好偏差(SPB)会严重损害该方法的可扩展性和可信度,表现为大模型在评价时系统性地偏好或贬低自身生成内容。现有测量方法依赖昂贵的人工标注,且混淆生成能力与评价立场,难以大规模应用。为此,我们提出一个完全自动化的框架,通过构造质量差异极小的响应对,实现无需人工标准即可统计分离判别能力与偏见倾向。对20个主流大模型的实证分析表明,先进能力往往与低偏见无关,甚至呈负相关。为缓解此问题,我们提出基于认知负荷分解的结构化多维度评估策略,在平均上将自偏好偏差降低31.5%。

原文摘要 · Abstract (English)

LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard construction, quality control, and so on. However, the scalability and trustworthiness of this approach can be substantially distorted by Self-Preference Bias (SPB), which is a directional evaluative deviation in which LLMs systematically favor or disfavor their own generated outputs during evaluation. Existing measurements rely on costly human annotations and conflate generative capability with evaluative stance, and thus are impractical for large-scale deployment in real-world systems. To address this issue, we introduce a fully automated framework to quantifying and mitigating SPB, which constructs equal-quality pairs of responses with negligible quality differences, enabling statistical disentanglement of discriminability from bias propensity without human gold standards. Empirical analysis across 20 mainstream LLMs reveals that advanced capabilities are often uncorrelated, or even negatively correlated, with low SPB. To mitigate this bias, we propose a structured multi-dimensional evaluation strategy grounded in cognitive load decomposition, which reduces SPB by 31.5\% on average.

大模型评估自偏好偏差自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。