通过多维度上下文增强,显著提升Transformer模型的作文自动评分效果。
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
- 引入多种上下文特征,改进Transformer在作文评分中的表现。
- 综合上下文后模型均值Kappa达0.823,部分数据集达0.8697。
- 方法通用性强,可适配任意作文评分模型,适合教育自动化研究者。
自动化作文评分(AES)因教育自动化需求而兴起,提供客观且低成本的长文本评估方案。尽管已有大量研究,但近年发现其他深度学习架构优于基于Transformer的模型。尽管Transformer在众多任务中表现优异,这一差距促使我们通过上下文增强来改进其在AES中的性能。本研究基于ASAP-AES数据集,分析多种上下文因素对Transformer模型的影响。最优模型融合多维上下文,在整个数据集上实现均值加权肯德尔协调系数(Quadratic Weighted Kappa)0.823,针对单个作文集训练时达到0.8697。该方法显著超越此前基于Transformer的模型,仅比当前最优深度学习模型(按作文集训练)平均低3.83%,且在八组数据集中有三组表现更优。该增强策略与模型架构无关,可无缝集成至任意AES模型。因此,上下文增强是一种通用、高效的提升自动评分能力的方法,有助于推动教育评价的智能化发展。
原文摘要 · Abstract (English)
Automated Essay Scoring (AES) has emerged to prominence in response to the growing demand for educational automation. Providing an objective and cost-effective solution, AES standardises the assessment of extended responses. Although substantial research has been conducted in this domain, recent investigations reveal that alternative deep-learning architectures outperform transformer-based models. Despite the successful dominance in the performance of the transformer architectures across various other tasks, this discrepancy has prompted a need to enrich transformer-based AES models through contextual enrichment. This study delves into diverse contextual factors using the ASAP-AES dataset, analysing their impact on transformer-based model performance. Our most effective model, augmented with multiple contextual dimensions, achieves a mean Quadratic Weighted Kappa score of 0.823 across the entire essay dataset and 0.8697 when trained on individual essay sets. Evidently surpassing prior transformer-based models, this augmented approach only underperforms relative to the state-of-the-art deep learning model trained essay-set-wise by an average of 3.83\% while exhibiting superior performance in three of the eight sets. Importantly, this enhancement is orthogonal to architecture-based advancements and seamlessly adaptable to any AES model. Consequently, this contextual augmentation methodology presents a versatile technique for refining AES capabilities, contributing to automated grading and evaluation evolution in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。