arXiv:2502.08450cs.CLcs.AI2025-02NAACL被引 2

让作文评分模型更懂语法,跨题目也能准判

Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay Scoring

  • 用语法纠错信息辅助训练,学通用作文特征
  • 跨题目评分准确率显著提升,语法相关维度改进明显
  • 适合需要泛化能力的自动阅卷场景

在自动作文评分(AES)中,近期研究转向跨题目设置,以应对实际应用需求。然而,以往方法依赖特定题目的作文-分数对,难以获得提示无关的作文表征。本文提出语法感知的跨提示特质评分模型(GAPS),通过语法纠错技术获取作文中的语法修正信息,并设计模型融合原始与修正后的文本,使模型在训练时聚焦于通用特征。实验证明,该方法在提示无关和语法相关特质上均有显著提升。尤其在最具挑战性的跨提示场景下,整体评分一致性(QWK)显著提高,凸显其对未见题目的评估优势。

原文摘要 · Abstract (English)

In automated essay scoring (AES), recent efforts have shifted toward cross-prompt settings that score essays on unseen prompts for practical applicability. However, prior methods trained with essay-score pairs of specific prompts pose challenges in obtaining prompt-generalized essay representation. In this work, we propose a grammar-aware cross-prompt trait scoring (GAPS), which internally captures prompt-independent syntactic aspects to learn generic essay representation. We acquire grammatical error-corrected information in essays via the grammar error correction technique and design the AES model to seamlessly integrate such information. By internally referring to both the corrected and the original essays, the model can focus on generic features during training. Empirical experiments validate our method's generalizability, showing remarkable improvements in prompt-independent and grammar-related traits. Furthermore, GAPS achieves notable QWK gains in the most challenging cross-prompt scenario, highlighting its strength in evaluating unseen prompts.

自动评分语法纠错跨提示作文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。