LLM作文评分存在系统性偏差,需小样本校正后才能可靠使用。
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
- 用关键词提示词比长篇评分标准更有效提升多维度评分准确率
- 模型对语法等低阶特征评分偏苛刻,且偏差稳定可测
- 仅需少量人工标注数据即可发现并校正关键偏差,适合教育评估落地
尽管大语言模型在教育评估中应用日益广泛,但其与人类评分的一致性仍不明确。我们系统评估了指令微调的开放权重模型在三个公开作文评分数据集(ASAP 2.0、ELLIPSE、DREsS)上的表现,涵盖整体评分与分项评分。结果表明,强模型在整体评分上与人类共识得分达到中高一致(加权二次卡帕系数约0.6),但该一致性无法推广至分项评分。尤其在低阶关注(LOC)特征如语法和标点上,模型呈现显著且稳定的负向偏差,即评分普遍低于人类。此外,简洁的关键词提示词在多维度评分中优于冗长的评分标准提示词。通过计算95%置信区间排除零偏所需的最小样本量,发现LOC偏差通常在极小验证集下即可检测,而高阶关注(HOC)特征则需更大样本。研究支持先校正偏差再部署的策略:利用小规模人工标注集估计并修正系统性评分偏移,无需大规模微调。
原文摘要 · Abstract (English)
Despite growing interest in using Large Language Models (LLMs) for educational assessment, it remains unclear how closely they align with human scoring. We present a systematic evaluation of instruction-tuned LLMs across three open essay-scoring datasets (ASAP 2.0, ELLIPSE, and DREsS) that cover both holistic and analytic scoring. We analyze agreement with human consensus scores, directional bias, and the stability of bias estimates. Our results show that strong open-weight models achieve moderate to high agreement with humans on holistic scoring (Quadratic Weighted Kappa about 0.6), but this does not transfer uniformly to analytic scoring. In particular, we observe large and stable negative directional bias on Lower-Order Concern (LOC) traits, such as Grammar and Conventions, meaning that models often score these traits more harshly than human raters. We also find that concise keyword-based prompts generally outperform longer rubric-style prompts in multi-trait analytic scoring. To quantify the amount of data needed to detect these systematic deviations, we compute the minimum sample size at which a 95% bootstrap confidence interval for the mean bias excludes zero. This analysis shows that LOC bias is often detectable with very small validation sets, whereas Higher-Order Concern (HOC) traits typically require much larger samples. These findings support a bias-correction-first deployment strategy: instead of relying on raw zero-shot scores, systematic score offsets can be estimated and corrected using small human-labeled bias-estimation sets, without requiring large-scale fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。