用学习者写作数据继续预训练,能提升英语水平测试的自动评分效果。
Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

- 在学习者作文数据上继续预训练,让模型更懂二语写作特点。
- 针对B1-B2水平数据微调,对FCE考试评分提升更稳定。
- 但跨数据集迁移效果不一致,需匹配评估场景才有效。
近期自动作文评分(AES)研究多采用通用英文预训练的Transformer模型,但这类模型可能无法充分代表第二语言学习者的写作特征。本研究探究在EFCAMDAT学习者语料上进行领域自适应持续预训练(DAPT)是否能提升Transformer模型在英语能力测试中的表现。我们在FCE和IELTS两个数据集上,对三种Transformer编码器进行DAPT,并评估其在本域评分与少量样本跨数据集迁移的表现。全语料DAPT在不同模型、数据集和指标下结果混杂。进一步分析表明,这些不一致结果部分源于EFCAMDAT与下游数据集在水平、文体和交际目的上的差异。基于CEFR分级的消融实验显示,使用对齐水平子集进行针对性DAPT比全语料更可靠地提升下游评分,尤其在包含B1-B2水平的FCE上。然而,这种收益并未显著改善跨数据集迁移。总体而言,当预训练数据与下游评估设置足够匹配时,学习者语料的持续预训练可提升本域自动评分;但不能自动增强跨数据集泛化能力。
原文摘要 · Abstract (English)
Recent automated essay scoring (AES) studies increasingly use pretrained transformer models, but these models are usually pretrained on general-domain English and may under-represent second-language learner writing. This study investigates whether domain-adaptive continued pretraining (DAPT) on the EFCAMDAT learner corpus improves transformer-based AES for English proficiency tests. We apply DAPT to three transformer encoders and evaluate them on FCE and IELTS in both in-domain scoring and few-shot cross-dataset transfer. Full-corpus DAPT produces mixed results across models, datasets, and metrics. Further analyses suggest that these mixed effects are partly explained by mismatches in proficiency, genre, and communicative purpose between EFCAMDAT and the downstream datasets. A proficiency-based ablation shows that targeted DAPT using CEFR-aligned subsets improves downstream scoring more reliably than full-corpus DAPT, especially for FCE with B1--B2 data. However, these gains do not consistently improve cross-dataset transfer. Overall, the findings suggest that continued pretraining on a learner-writing corpus can benefit in-domain AES for English assessment when the pretraining data is sufficiently aligned with the downstream assessment settings. However, it does not automatically improve transferability across different English proficiency test datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。