arXiv:2602.02219cs.CL2026-02被引 5

发现大模型评分时存在位置偏好,影响评测公正性。

Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge

  • 将评分视为多选题,揭示模型对选项位置的依赖。
  • 不同模型偏好首项或末项,且评分顺序也会影响结果。
  • 随机打乱评分项顺序可有效降低偏差,尤其对强偏模型。

大语言模型作为评分者(LLM-as-a-judge)在基于量规的评估中广泛应用,但其隐含的多选题结构导致位置偏差:模型倾向于选择量规列表中特定位置的评分选项。通过跨多个模型和数据集的受控实验,我们发现该偏差具有一致性,但方向因模型而异——部分模型偏好首位,部分偏好末位。此外,当提示同时满足多个标准时,标准的排序本身也会改变最终得分。我们进一步测试了打乱量规选项顺序的方法,发现虽无法完全消除偏差,但仅需少量随机排列即可显著提升偏差严重模型的评分与人工标注的相关性。结果表明,基于量规的评估本质上是带有可测量、模型特异性位置偏差的多选任务,且少量随机化即可有效缓解多数模型的误差。

原文摘要 · Abstract (English)

Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examined point-wise or pair-wise evaluation protocols; in contrast, our focus is on rubric-based evaluation, which has been attracting increasing attention owing to its utility for training models in domains where verification is otherwise difficult. In this work, we show that rubric-based evaluation implicitly resembles a multiple-choice setting and therefore exhibits position bias: LLMs tend to prefer score options that appear at specific positions within the rubric list. Through controlled experiments across multiple models and datasets, we demonstrate that this position bias is consistent. Its direction, however, is model-specific: some judges favor the first option, while others favor the last. We further identify a second, orthogonal axis of bias: when a prompt scores several criteria simultaneously, the ordering of the criteria itself shifts the resulting scores. We additionally explore permuting the order of the rubric options as a means of mitigating position bias, and find that although the bias can be attenuated, improvements in the correlation between model judgments and human annotations are obtained primarily for models that exhibit strong bias. Our results recast rubric-based LLM-as-a-judge as a multiple-choice problem with measurable, model-specific position bias, and we further confirm that only a small number of random order permutations are sufficient to reduce the error introduced by this bias for the majority of models.

大模型评测位置偏差量规评分可信评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。