让语法纠错与解释相互增强,提升学习者可理解性。
Corrections Meet Explanations: A Unified Framework for Explainable Grammatical Error Correction
- 将纠错与解释任务统一生成,互相促进
- 在2万样本数据集上表现优于单任务模型
- 提供去噪版数据集,提升训练评估可靠性
语法错误纠正(GEC)在面向语言学习者时面临可解释性挑战。现有研究多聚焦于预先提取的错误解释,忽视了解释与纠正之间的关联。为此,我们提出EXGEC框架,以生成方式统一整合解释与纠正任务,主张两者相互促进。实验基于最新人工标注的可解释性GEC数据集EXPECT(约2万样本),并发现其存在显著噪声,可能影响模型训练与评估。因此,我们构建了去噪版本EXPECT-denoised,确保更客观的训练与评估框架。在BART、T5和Llama3等多种NLP模型上,EXGEC模型在两项任务中均超越单任务基线,验证了该方法的有效性。
原文摘要 · Abstract (English)
Grammatical Error Correction (GEC) faces a critical challenge concerning explainability, notably when GEC systems are designed for language learners. Existing research predominantly focuses on explaining grammatical errors extracted in advance, thus neglecting the relationship between explanations and corrections. To address this gap, we introduce EXGEC, a unified explainable GEC framework that integrates explanation and correction tasks in a generative manner, advocating that these tasks mutually reinforce each other. Experiments have been conducted on EXPECT, a recent human-labeled dataset for explainable GEC, comprising around 20k samples. Moreover, we detect significant noise within EXPECT, potentially compromising model training and evaluation. Therefore, we introduce an alternative dataset named EXPECT-denoised, ensuring a more objective framework for training and evaluation. Results on various NLP models (BART, T5, and Llama3) show that EXGEC models surpass single-task baselines in both tasks, demonstrating the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。