arXiv:2506.22316cs.CL2025-06中稿 · DASFAA 2026被引 50

首次揭示大模型评分裁判的三种隐藏偏见,提升自动评估可靠性

Evaluating Scoring Bias in LLM-as-a-Judge

  • 从评分提示本身出发,识别出评语顺序、分数标识、参考答案等三类新偏见
  • 实验发现顶尖大模型在评分中仍存在显著偏差,影响评估公平性
  • 提供量化框架与数据生成工具,适合评估系统设计者使用

将大语言模型作为自动评判者(LLM-as-a-Judge)已成为大模型发展中的关键范式,可为复杂任务提供可扩展的反馈。然而,这类评判者的可靠性受多种偏见影响。现有研究主要关注比较式评估中的偏见,而更贴近工业应用的评分式评估——即给出绝对分值——却鲜有深入探讨。为此,本文首次专门考察评分偏见问题,将焦点从评价目标转向评分提示本身。我们正式定义评分偏见,并识别出三种此前未被研究的新类型:评分标准顺序偏见、分数标识偏见和参考答案分数偏见。我们提出一个综合性量化框架,包含多维度度量指标与自动化数据合成流程,构建专用评估语料库。实验表明,即使是当前最先进的大模型也严重受到这些偏见影响。分析结果为设计更稳健的评分提示、缓解新发现的偏见提供了可操作的洞见。

原文摘要 · Abstract (English)

The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the reliability of these judges is compromised by various biases. Existing research has heavily concentrated on biases in comparative evaluations. In contrast, scoring-based evaluations-which assign an absolute score and are often more practical in industrial applications-remain under-investigated. To address this gap, we undertake the first dedicated examination of scoring bias in LLM judges. We shift the focus from biases tied to the evaluation targets to those originating from the scoring prompt itself. We formally define scoring bias and identify three novel, previously unstudied types: rubric order bias, score ID bias, and reference answer score bias. We propose a comprehensive framework to quantify these biases, featuring a suite of multi-faceted metrics and an automatic data synthesis pipeline to create a tailored evaluation corpus. Our experiments empirically demonstrate that even the most advanced LLMs suffer from these substantial scoring biases. Our analysis yields actionable insights for designing more robust scoring prompts and mitigating these newly identified biases.

大模型评估评分偏见自动化评测提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。