arXiv:2507.03772cs.LGstat.ML2025-07被引 8

用统计模型检测自动评分器的评分偏差,提升评估可靠性。

Skewed Score: A statistical framework to assess autograders

  • 基于贝叶斯广义线性模型,统一建模评分者与被评内容特征。
  • 可量化不同评分者间的差异,识别系统性偏差来源。
  • 适合关注LLM评估公平性与结果可信度的研究者。

大语言模型(LLM)输出的评估越来越多地由其他LLM完成,即“LLM作为评判者”(LLM-as-a-judge),又称自动评分器。尽管自动评分器具备可扩展性,但其可靠性参差不齐,可能因响应类型、评分方法或领域特定性等因素表现出系统性偏差。本文提出一种基于贝叶斯广义线性模型(GLMs)的统计框架,使研究人员能在解决核心研究问题(如LLM评估)的同时,同步评估自动评分器的表现。该方法将评分结果(如分数或成对偏好)建模为评分者属性(如人类与自动评分器)和被评项特征(如响应长度或生成模型)的函数,从而显式量化评分差异与潜在偏差。此外,该方法可增强传统一致性指标(如评分者间一致率),提供不确定性估计并澄清分歧来源。整体上,该框架提升了自动评分器在LLM评估中的鲁棒性与可解释性,支持性能分析与偏见检测。

原文摘要 · Abstract (English)

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they have shown mixed reliability and may exhibit systematic biases, depending on response type, scoring methodology, domain specificity, or other factors. Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to simultaneously assess their autograders while addressing their primary research questions (e.g., LLM evaluation). Our approach models evaluation outcomes (e.g., scores or pairwise preferences) as a function of properties of the grader (e.g., human vs. autograder) and the evaluated item (e.g., response length or the LLM that generated it), allowing for explicit quantification of scoring differences and potential biases within a unified framework. In addition, our method can be used to augment traditional metrics such as inter-rater agreement, by providing uncertainty estimates and clarifying sources of disagreement. Overall, this approach contributes to more robust and interpretable use of autograders in LLM evaluation, enabling both performance analysis and bias detection.

自动评分统计建模偏差检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。