提出一种更公平的语法纠错评估框架,支持多参考答案。
A Formal Framework for Fluency-based Multi-Reference Evaluation in Grammatical Error Correction
- 用n-gram相似度聚合多个正确修正结果,避免依赖单一标准答案。
- 四种聚合策略在4种语言数据上均有效捕捉流畅性与覆盖度。
- 适合多语言、生成式纠错任务,能包容合理语法差异。
语法纠错评估需反映人类修正的多样性,而非偏倚单一参考答案。现有方法多基于编辑且以英语为主,依赖系统与参考答案间严格对齐,限制了其在多语言和生成式场景的应用。本文提出一种形式化框架,用于基于流畅性的多参考评估,将n-gram相似度视为对多个合法修正结果的聚合问题。在此框架下,我们通过四种聚合策略(选择最优、简单平均、加权平均、合并计数)实例化GLEU,并分析其有界性、单调性及对参考答案变化的敏感性。在捷克语、爱沙尼亚语、乌克兰语和中文语料上的实验表明,这些策略捕捉了流畅性与覆盖度的不同侧面。该框架将多参考评估统一为一个原则性强、以流畅性为导向的方法,兼顾语言多样性而不惩罚合理变异。
原文摘要 · Abstract (English)
Evaluating grammatical error correction requires metrics that reflect the diversity of valid human corrections rather than privileging a single reference. Existing frameworks, largely edit-based and English-centric, rely on rigid alignments between system and reference edits, limiting their applicability in multilingual and generative settings. This paper introduces a formal framework for \textit{fluency-based multi-reference evaluation}, framing $n$-gram similarity as an aggregation problem over multiple legitimate corrections. Within this formulation, we instantiate GLEU through four aggregation strategies--\textsc{select-best}, \textsc{simple-average}, \textsc{weighted-average}, and \textsc{merged-counts}--and analyze their properties of boundedness, monotonicity, and sensitivity to reference variation. Empirical results on Czech, Estonian, Ukrainian, and Chinese corpora show that these strategies capture complementary aspects of fluency and coverage. The framework unifies multi-reference evaluation into a principled, fluency-oriented approach that incorporates linguistic diversity without penalizing legitimate variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。