自动评估语法纠错中每处修改的重要性,提升多语言纠错评估效率。
Scoring Edit Impact in Grammatical Error Correction via Embedded Association Graphs

- 构建嵌入关联图捕捉编辑间的潜在依赖关系。
- 通过困惑度评分量化每处修改对句子流畅度的贡献。
- 跨4语种、4数据集表现稳定,适合多语言纠错系统评估。
语法错误纠正(GEC)系统会生成一系列修改来修正错误句子。传统评估依赖人工标注,但同一句子可能存在多种有效修正方式,现有评估难以适配多样化应用场景。近期元评估方法依赖多人判断多个参考答案,难以扩展至大规模数据。本文提出新任务:在GEC中评分编辑影响,旨在自动估计每个修改的重要性。为此,我们引入基于嵌入关联图的评分框架,该图捕捉编辑间的隐含依赖及句法相关性,将编辑分组为连贯单元,并通过困惑度评分估算每处修改对句子流畅度的贡献。在4个GEC数据集、4种语言和4个GEC系统上的实验表明,该方法持续优于多种基线。进一步分析显示,嵌入关联图能有效捕获跨语言的编辑结构依赖关系。
原文摘要 · Abstract (English)
A Grammatical Error Correction (GEC) system produces a sequence of edits to correct an erroneous sentence. The quality of these edits is typically evaluated against human annotations. However, a sentence may admit multiple valid corrections, and existing evaluation settings do not fully accommodate diverse application scenarios. Recent meta-evaluation approaches rely on human judgments across multiple references, but they are difficult to scale to large datasets. In this paper, we propose a new task, Scoring Edit Impact in GEC, which aims to automatically estimate the importance of edits produced by a GEC system. To address this task, we introduce a scoring framework based on an embedded association graph. The graph captures latent dependencies among edits and syntactically related edits, grouping them into coherent groups. We then perform perplexity-based scoring to estimate each edit's contribution to sentence fluency. Experiments across 4 GEC datasets, 4 languages, and 4 GEC systems demonstrate that our method consistently outperforms a range of baselines. Further analysis shows that the embedded association graph effectively captures cross-linguistic structural dependencies among edits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。