arXiv:2606.02009cs.CL2026-06Transactions of th…

评估法语作文自动评分模型的可靠性与公平性,提升评分系统可信度。

Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French

论文配图:Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French
图 1 · 摘自论文原文
  • 基于改进的论证验证框架,多维度评估模型性能
  • 在2.7万篇法语作文上对比8种模型架构,泛化集含961篇
  • 关注评分一致性、语言特征关联及与人工评分的误差差异

在自动作文评分(AES)领域,基准测试实践导致评估方式趋于简化,与更全面的评价框架(如基于论证的验证框架,ABV)建议相悖。本文提出一种增强且更实用的ABV框架,包含公平性分析、与语言特征的相关性、预测误差评估以及模型与人工评分者的一致性比较。将该框架应用于法语AES任务,对8种模型架构在包含27,000篇作文(每篇2名评分者)的考试数据集和961篇作文(至少9名评分者)的泛化数据集上进行评估。结果表明,采用该框架能更深入理解AES模型的能力与局限,同时推动了法语自动评分技术的前沿水平。

原文摘要 · Abstract (English)

In Automated Essay Scoring (AES), benchmarking practices have fostered minimalist evaluation practices, in contrast with the broader-view recommendations of evaluation frameworks, such as the argument-based validation framework (ABV), which argued in favor of a multidimensional assessment of systems, especially in the context of high-stakes language tests. In this paper, we introduce an enhanced and more practical version of the ABV framework, incorporating fairness analysis, correlations with linguistic features, prediction error evaluation, and model agreement compared with human raters. Applying this framework to French AES, we compare 8 model architectures on a corpus of 27k exam essays (2 raters each) and a generalization corpus of 961 essays (at least nine raters each). Our analyses illustrate the benefits of applying the ABV framework to better understand the capabilities and pitfalls of AES models, while also advancing the state-of-the-art for French AES.

自动评分法语评估框架模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。