arXiv:2608.29517cs.CL2026-08

检验大模型作文评分的偏差与稳定性,发现其评分差异显著但光环效应不超人类水平。

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

  • 将大模型视为评分员,采用多面拉斯克模型等方法系统检测评分偏差
  • 模型间评分差异达219分(满分1000),版本更新导致最大133分偏移
  • 实证显示模型光环效应未超过训练人类评分员范围,适合教育评估研究者参考

大型语言模型在学习分析中被越来越多用作作文评分工具,但评价几乎仅依赖一致性统计。教育测量学指出评分员存在严重程度差异、光环效应及随时间漂移。本文将大模型视为评分员,在双语公开数据集(ENEM/Essay-BR;ASAP)上开展预注册的评分员效应测试(多面拉斯克严重度、残差光环、信度/决策研究、跨版本变化、差异功能)。共分析2,377篇作文,12名模型评分员,4家提供方,5种版本对比,结果以评分张量形式公开。在ENEM上,模型评分严重度跨度达219分(满分1000);在ASAP上,评分员群体分布占总分范围的15%-33%,远超训练人类间差距(约1%)。模型与人类评分相关性仅为.47-.56,处于低区分度区间。所有五种版本对比均使评分严重度发生显著变化(超越家族校正置换检验阈值,最高达133分),其中一名模型因身份异常被提前淘汰。两项预注册检验返回真实零假设:经严重度调整后排名反转未通过置换检验,且“无声漂移”被证伪——在五组对比中有四组,一致性随严重度变化而变动。重复实验显示自一致性良好(k≤2时φ≥.80),但未达人类准确度水平。同一仪器检查推翻了先前光环效应比较:在相同工具与校准条件下,未发现模型光环效应显著超过训练人类范围。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and "silent drift" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi>=.80 at k<=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.

大模型评估评分偏差教育测量可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。