arXiv:2608.18091cs.CLcs.AI2026-08

LLM裁判会因自我标签而偏袒自己,即使不看内容也这样。

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

论文配图:Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
图 1 · 摘自论文原文
  • 用叙事选择任务分离风格与质量,避免混淆
  • 盲评下自偏消失,但标签本身仍能引发双向偏差
  • 适合研究AI评估可靠性或模型偏见的学者

随着大模型作为裁判系统日益普及,模型偏好自身输出的现象——即自我偏好——引发了对评估可靠性的担忧。然而,现有研究主要集中在生成文本上,导致风格特征与回答质量混杂,无法区分真实的自我偏好与其他干扰因素。为此,我们改变评估对象:让十名大模型评估叙事约束选择,这些选择不带有模型特有的风格痕迹,但仍保留可识别的模型专属签名。通过两个实验,我们得到不同发现:在盲评条件下,控制选择质量与评判严格度后,自我偏好基本消失,三类评分维度上无偏,第四类反而认为自己的选择原创性更低;但在质量匹配情况下,仅凭自我或他人标签(不暴露模型身份)即可引发双向偏差:大模型裁判会提高对自己标签的选择评分,降低对他人标签的评分,无论实际来源。本研究贡献两点:1)作者归属识别是评估偏差的独立驱动因素;2)开放、无需真实答案的任务可作为研究大模型裁判行为的受控工具。

原文摘要 · Abstract (English)

As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.

大模型评估偏见分析自回归偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。