arXiv:2601.22548cs.CLcs.AI2026-01被引 4

发现大模型评自分身输出偏爱,实为评估质量差异所致。

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

  • 通过对比模型自评与他评的投票分布,分离出评估偏差来源。
  • 仅51%的自我偏好结果在质量基准下仍显著,覆盖89.6%概率质量。
  • 揭示评估不确定性导致的重叠,推动更严谨的评测文档规范。

近期研究发现,大语言模型在充当评判者时倾向于偏好自身生成内容,威胁自动化后训练与评估流程的可靠性。然而,难以区分这种倾向是由自恋行为还是实验混杂因素导致。具体而言,当模型评估自己无法解答的问题时,其自评结果可能并非源于作者身份,而是由评估者自身质量决定。本文通过直接比较模型在自评与他评场景下的投票分布,建立评估质量基线。结果显示,此前发现中仅有51%的例子在该基线检验下仍具统计显著性,但覆盖了89.6%的总自我偏好概率质量。此外,通过分析投票分布熵值,暗示存在由不确定性引起的判断重叠现象。本方法有助于更谨慎地记录与审查评判偏差问题。

原文摘要 · Abstract (English)

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workflows. However, it is difficult to disentangle which behaviors are explained by narcissism versus experimental confounds. Specifically, LLM evaluators may deliver self-preferring verdicts when comparing responses to questions they fail on; these verdicts may not depend on the identity of the author, but on evaluator quality. We correct this by directly comparing the judge's voting distribution in cases where it evaluates itself versus another model. This evaluator quality baseline reveals that only 51% of examples in previous findings retain statistical significance against this null hypothesis, covering 89.6% of total self-preference probability mass. Finally, we compare the entropy of voting distributions, suggesting uncertainty-driven overlap, and show that our procedure enables more careful documentation against the backdrop of judge-bias research.

大模型评估自评偏差评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。