arXiv:2604.19131cs.AI2026-04中稿 · publication at AIE…被引 2

提出两种新天花板指标,评估自动作文评分的极限与实用目标

Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory

  • 基于经典测试理论推导出理论与人类水平双天花板
  • 实证显示现有模型距理想性能仍有显著提升空间
  • 适合评估和优化自动评分系统,尤其关注部署可行性

自动作文评分(AES)通常在公开基准上以二次加权卡帕系数(QWK)衡量。然而,由于评分标签由人工评定且不可避免存在误差,当前尚不清楚理论上可达到的最高QWK值,以及实际部署所需的足够性能水平。本文基于经典测试理论中的信度概念,从标准双评者基准中推导出两个数据集相关的QWK天花板:一是理想模型在标签噪声下的理论上限,即完美预测潜在真实分数时能达到的最大QWK;二是人类水平天花板,代表与人类评分误差相当的模型所能达到的实用目标,适用于替代单个评分员的场景。研究还表明,常被用作参考的真人对真人QWK可能低估真实天花板。模拟实验验证了该方法的有效性,真实基准实验则揭示了现代AES模型的当前表现与剩余提升空间。

原文摘要 · Abstract (English)

Automated essay scoring (AES) is commonly evaluated on public benchmarks using quadratic weighted kappa (QWK). However, because benchmark labels are assigned by human raters and inevitably contain scoring errors, it remains unclear both what QWK is theoretically attainable and what level is practically sufficient for deployment. We therefore derive two dataset-specific QWK ceilings based on the reliability concept in classical test theory, which can be estimated from standard two-rater benchmarks without additional annotation. The first is the theoretical ceiling: the maximum QWK that an ideal AES model that perfectly predicts latent true scores can achieve under label noise. The second is the human-like ceiling: the QWK attainable by an AES model with human-level scoring error, providing a practical target when AES is intended to replace a single human rater. We further show that human--human QWK, often used as a ceiling reference, can underestimate the true ceiling. Simulation experiments validate the proposed ceilings, and experiments on real benchmarks illustrate how they clarify the current performance and remaining headroom of modern AES models.

自动评分评估基准信度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。