arXiv:2505.23315cs.CLcs.AI2025-05ACL被引 6

通过有序置信度建模,提升作文评分系统可靠性。

Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments

  • 将置信度预测转化为有序分类任务,利用CEFR等级结构
  • 新损失函数使模型在47%分数上实现100%等级一致
  • 适合对评分可靠性要求高的教育评估场景

自动作文评分(AES)面临的重要伦理挑战是确保评分仅在高可靠性时才发布。置信度建模通过为每个自动评分分配置信度分数来应对这一问题。本文将置信度估计视为分类任务:判断一个AES生成的分数是否准确地将考生归入相应的CEFR等级。尽管这是二分类问题,我们通过两种方式利用评分领域的内在粒度:首先,采用分数分箱将任务重构为n元分类;其次,引入一种新型核加权有序类别交叉熵(KWOCCE)损失函数,融入CEFR标签的有序结构。最佳模型达到0.97的F1分数,使系统能释放47%的评分,实现100%的CEFR一致性,99%的评分至少达到95%的一致性——相比之下,独立使用AES模型并全部释放评分时,一致性约为92%。

原文摘要 · Abstract (English)

A key ethical challenge in Automated Essay Scoring (AES) is ensuring that scores are only released when they meet high reliability standards. Confidence modelling addresses this by assigning a reliability estimate measure, in the form of a confidence score, to each automated score. In this study, we frame confidence estimation as a classification task: predicting whether an AES-generated score correctly places a candidate in the appropriate CEFR level. While this is a binary decision, we leverage the inherent granularity of the scoring domain in two ways. First, we reformulate the task as an n-ary classification problem using score binning. Second, we introduce a set of novel Kernel Weighted Ordinal Categorical Cross Entropy (KWOCCE) loss functions that incorporate the ordinal structure of CEFR labels. Our best-performing model achieves an F1 score of 0.97, and enables the system to release 47% of scores with 100% CEFR agreement and 99% with at least 95% CEFR agreement -compared to approximately 92% (approx.) CEFR agreement from the standalone AES model where we release all AM predicted scores.

自动评分置信度教育评估有序分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。