arXiv:2601.03444cs.CLcs.AI2026-01被引 16

0-5分制让AI评分最接近人类判断。

Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale

  • 对比0-5、1-10等评分尺度过,0-5分制下人类与AI评分最一致。
  • 在主观题上,不同评分尺度过会导致人类与AI评分差异高达23%。
  • 该研究提示评估系统设计需关注评分尺度,尤其要检查性别等子群体差异。

大型语言模型(LLMs)正被广泛用作自动化评价者,但已有研究表明,当提示语改变时,这些模型的评分常缺乏一致性。然而,评分尺度本身的影响仍鲜有研究。本研究通过比较人类与LLM两类评价者,在三种评分尺度下对六个基准任务(包含客观题、开放式主观题及混合任务)进行评分。使用组内相关系数(ICC)衡量绝对一致性,发现对于主观任务,不同评分尺度下LLM评分的一致性存在显著差异,且评分尺度会大幅影响人类与LLM之间的评分一致性,即便组内可靠性本身较高。综合所有任务后,0-5分制下的人类-模型对齐程度最高。研究还发现,合并可靠性可能掩盖不同基准间的异质性,并揭示出在性别子群体中存在系统性对齐差异,强调了评分尺度设计和子群体诊断在构建可靠LLM评价协议中的关键作用。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is altered. However, the effect of the grading scale itself remains underexplored. We study the LLM-as-a-judge problem by comparing two kinds of raters: humans and LLMs. We collect ratings from both groups on three scales and across six benchmarks that include objective, open-ended subjective, and mixed tasks. Using intraclass correlation coefficients (ICC) to measure absolute agreement, we find that LLM judgments are not perfectly consistent across scales on subjective benchmarks, and that the choice of scale substantially shifts human-LLM agreement, even when within-group panel reliability is high. Aggregated over tasks, the grading scale of 0-5 yields the strongest human-LLM alignment. We further demonstrate that pooled reliability can mask benchmark heterogeneity and reveal systematic subgroup differences in alignment across gender groups, strengthening the importance of scale design and sub-level diagnostics as essential components of LLM-as-a-judge protocols.

AI评价评分尺度人类对齐模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。