arXiv:2504.08537cs.CL2025-04被引 1

用词块频率提升英语写作自动评分准确度,尤其帮助区分初学者和中等水平写作者。

Lexical Bundle Frequency as a Construct-Relevant Candidate Feature in Automated Scoring of L2 Academic Writing

  • 将词块出现频率作为特征加入评分模型,增强语言层面的判断能力。
  • 模型与人工评分一致性提升5.63%,低/中水平作文准确率显著改善。
  • 非提示性词块频率对水平判断更敏感,适合评估发展中的二语写作者。

自动化评分系统在评估二语学术写作中日益普及,但需持续优化以确保构念效度。尽管先前研究指出词块(频繁出现的多词组合)可辅助评估,其在评分模型中的实证整合仍需深入探索。本研究检验了将词块频率特征纳入托福独立写作任务评分模型的效果。分析来自TOEFL11语料库的抽样子集(N=1,225篇作文,9种母语背景),由ETS培训评卷人评定为低、中、高三个等级。提取3至9个词的词块,区分提示相关与非提示类型。对比基于传统语言特征(如语法、衔接、复杂度)的基准支持向量机(SVM)模型与加入三项聚合词块频率特征(提示类总频次、非提示类总频次、总体总频次)的扩展模型。结果显示,词块频率(尤其是非提示类)与写作水平存在显著但效应量较小的关系(p < .05)。平均频率表明低水平作文整体使用更多词块。关键发现是,加入词块特征的模型与人工评分者的一致性提升(加权肯德尔系数+2.05%,总体肯德尔系数+5.63%),尤其在低水平(精确一致率+10.1%)和中等水平(肯德尔系数+14.3%)作文中表现突出。研究证明,整合聚合词块频率可助力开发更具语言学依据且更精准的自动化评分系统,尤其有助于区分发展中二语写作者。

原文摘要 · Abstract (English)

Automated scoring (AS) systems are increasingly used for evaluating L2 writing, but require ongoing refinement for construct validity. While prior work suggested lexical bundles (LBs) - recurrent multi-word sequences satisfying certain frequency criteria - could inform assessment, their empirical integration into AS models needs further investigation. This study tested the impact of incorporating LB frequency features into an AS model for TOEFL independent writing tasks. Analyzing a sampled subcorpus (N=1,225 essays, 9 L1s) from the TOEFL11 corpus, scored by ETS-trained raters (Low, Medium, High), 3- to 9-word LBs were extracted, distinguishing prompt-specific from non-prompt types. A baseline Support Vector Machine (SVM) scoring model using established linguistic features (e.g., mechanics, cohesion, sophistication) was compared against an extended model including three aggregate LB frequency features (total prompt, total non-prompt, overall total). Results revealed significant, though generally small-effect, relationships between LB frequency (especially non-prompt bundles) and proficiency (p < .05). Mean frequencies suggested lower proficiency essays used more LBs overall. Critically, the LB-enhanced model improved agreement with human raters (Quadratic Cohen's Kappa +2.05%, overall Cohen's Kappa +5.63%), with notable gains for low (+10.1% exact agreement) and medium (+14.3% Cohen's Kappa) proficiency essays. These findings demonstrate that integrating aggregate LB frequency offers potential for developing more linguistically informed and accurate AS systems, particularly for differentiating developing L2 writers.

自动评分词块分析二语写作语言特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。