arXiv:2602.01747cs.CLcs.LG2026-02被引 1

用三种新方法提升作文评分模型在少样本下的表现

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

  • 分两阶段微调+分数对齐+不确定度自训练,增强模型适应性
  • 32组数据下达91.2%全量数据性能,仅用约1000个标注样本
  • 适合资源有限但需高精度作文评分的教育场景

自动作文评分(AES)在教育中发挥重要作用,但真实场景中标签数据极度稀缺,制约了系统发展。本文提出三项技术提升AES在少样本与全量数据下的表现:首先采用低秩适配的两阶段微调策略,使模型更好适应目标题型作文;其次引入分数对齐技术,提升预测分数分布与真实分数的一致性;最后利用不确定性感知的自训练方法,通过伪标签扩充无标注数据,减少噪声传播。上述方法在DualBERT上实现,实验基于ASAP++数据集,并在TOEFL11和ELLIPSE上验证泛化能力。在ASAP++的32组数据设置下,三项技术协同可达到约1000个标注样本全量数据性能的91.2%。此外,分数对齐技术在少样本与全量设置下均持续提升性能,集成后在ASAP++全量设置下达到当前最优结果。

原文摘要 · Abstract (English)

Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world settings, the extreme scarcity of labeled data severely limits the development and practical adoption of robust AES systems. This study proposes a novel approach to enhance AES performance in both limited-data and full-data settings by introducing three key techniques. First, we introduce a Two-Stage fine-tuning strategy that leverages low-rank adaptations to better adapt an AES model to target prompt essays. Second, we introduce a Score Alignment technique to improve consistency between predicted and true score distributions. Third, we employ uncertainty-aware self-training using unlabeled data, effectively expanding the training set with pseudo-labeled samples while mitigating label noise propagation. We implement the above three key techniques on DualBERT. We conduct extensive experiments on the ASAP++ dataset, and additionally evaluate the proposed techniques on two other datasets, TOEFL11 and ELLIPSE, to examine their generalizability. In the 32-data setting on ASAP++, all three key techniques improve performance, and their integration achieves 91.2% of the full-data performance trained on approximately 1,000 labeled samples. In addition, the proposed Score Alignment technique consistently improves performance in both limited-data and full-data settings: e.g., it achieves state-of-the-art results in the full-data setting on ASAP++ when integrated into DualBERT.

自动评分少样本学习文本生成教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。