arXiv:2508.06709cs.CLcs.AI2025-08被引 47

提出统计方法,精准识别大模型自评偏见。

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

  • 通过对比模型自评与第三方评分差异,建模自偏见
  • 在5000+样本中发现GPT-4o等模型存在自评偏高
  • 揭示模型家族间也存在系统性偏袒现象

大型语言模型可作为快速可靠的评判者评估其他模型输出,但可能对自身输出给予过度偏好评分,即自偏见,这会扭曲真实性能评估。以往研究常将模型质量差异误判为偏见,或错误假设模型与人类评分分布一致。本文提出一种统计框架,明确形式化自偏见识别与估计的假设条件。该方法建模模型自评与其输出与其他模型输出之间的评分分布差异,同时考虑独立第三方评判者(如人类)提供的基础质量。该方法能在模型能力不同时仍可靠分离并量化自偏见,避免将真实性能差异误认为偏见。我们在包含超过5000个提示-完成对的大规模数据集上进行实证分析,涵盖专家人类标注及九个不同大模型裁判的评分。结果发现,GPT-4o和Claude 3.5 Sonnet等模型系统性地为其自身输出打更高分,且表现出家族偏见——对同家族其他模型输出也倾向于打更高分。研究揭示了使用大模型作评判者的潜在风险,并为解读自动化评估提供实用指导。

原文摘要 · Abstract (English)

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings to their own outputs, a phenomenon known as self-bias, which can distort evaluations of true model performance. Previous studies often conflate genuine differences in model quality with bias or incorrectly assume that evaluations from LLMs and humans follow the same rating distributions. In this work, we present a statistical framework that explicitly formalizes assumptions under which self-bias can be identified and estimated. Our method models the difference in the scoring distribution that LLM-as-a-judge assigns to its own completions compared to other models, while accounting for the underlying quality of the completions provided by an independent, third-party judge (e.g., humans). Our method reliably isolates and quantifies self-bias, even when models vary in ability, ensuring that genuine performance differences are not mistaken for self-bias. We conduct an empirical analysis of self-bias on a large dataset (>5000 prompt-completion pairs) consisting of expert human annotations and judgments from nine different LLM judges. We find that some models, such as GPT-4o and Claude 3.5 Sonnet, systematically assign higher scores to their own outputs. These models also display family-bias; systematically assigning higher ratings to outputs produced by other models of the same family. Our findings highlight potential pitfalls of using LLM judges and offer practical guidance to mitigate biases when interpreting automated evaluations.

大模型评估自偏见统计方法评测偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。