用英语语法分级体系分析二语语法,自动识别对错并评估水平。
Exploiting the English Grammar Profile for L2 grammatical analysis with LLMs
- 基于英语语法分级体系,用LLM检测学习者语法尝试
- 混合方法在语法水平评估中表现最佳,接近人工修正效果
- 关注成功尝试,支持正向反馈,适合语言教学研究
评估第二语言(L2)学习者的语法能力对于提供针对性反馈和衡量语言水平至关重要。本文提出一种新框架,利用英语语法分级体系(EGP),该体系将语法结构映射到欧洲共同语言参考框架(CEFR)的等级。通过分析学习者句子与修正句的配对,检测其对特定语法结构的尝试,并分类为成功或失败,从而实现细粒度反馈。同时,这些语法结构作为预测变量,用于自动化评估整体CEFR水平。比较了基于规则与基于LLM的分类器,发现LLM在需要语义和语用理解的复杂结构上表现更优,而规则方法在仅依赖形态或句法特征的结构上仍具竞争力。在水平评估中,结合规则预筛选与LLM的混合管道表现最强。此外,采用自动语法纠错的全自动流程也接近半自动系统性能,尤其在识别语法成功尝试方面表现优异。总体而言,该框架不仅关注错误,还强调学习者的成功尝试,有助于生成积极、建设性的反馈,为语法发展提供可操作的洞察。
原文摘要 · Abstract (English)
Evaluating the grammatical competence of second language (L2) learners is essential both for providing targeted feedback and for assessing proficiency. To achieve this, we propose a novel framework leveraging the English Grammar Profile (EGP), a taxonomy of grammatical constructs mapped to the proficiency levels of the Common European Framework of Reference (CEFR), to detect learners' attempts at grammatical constructs and classify them as successful or unsuccessful. This detection can then be used to provide fine-grained feedback. Moreover, the grammatical constructs are used as predictors of proficiency assessment by using automatically detected attempts as predictors of holistic CEFR proficiency. For the selection of grammatical constructs derived from the EGP, rule-based and LLM-based classifiers are compared. We show that LLMs outperform rule-based methods on semantically and pragmatically nuanced constructs, while rule-based approaches remain competitive for constructs that rely purely on morphological or syntactic features and do not require semantic interpretation. For proficiency assessment, we evaluate both rule-based and hybrid pipelines and show that a hybrid approach combining a rule-based pre-filter with an LLM consistently yields the strongest performance. Since our framework operates on pairs of original learner sentences and their corrected counterparts, we also evaluate a fully automated pipeline using automatic grammatical error correction. This pipeline closely approaches the performance of semi-automated systems based on manual corrections, particularly for the detection of successful attempts at grammatical constructs. Overall, our framework emphasises learners' successful attempts in addition to unsuccessful ones, enabling positive, formative feedback and providing actionable insights into grammatical development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。