arXiv:2506.13148cs.CLcs.AI2025-06中稿 · BEA-2025被引 10

改进大模型在最小修改语法纠错中的表现,提升准确性并修复数据集错误。

Adapting LLMs for Minimal-edit Grammatical Error Correction

  • 提出新训练策略,优化大模型在最小修改场景下的纠错能力。
  • 在BEA-test上达成单模型新最佳性能,显著提升纠错精度。
  • 修复常见数据集的标注错误,推动可复现研究和高质量训练。

仅使用解码器的大型语言模型在流畅性修正型英语语法纠错任务中表现出色,但在最小修改型英语语法纠错任务中的适配仍待深入探索。为提升其在最小修改方法中的有效性,本文探讨了错误率适应问题,并提出一种新型训练调度方法。实验表明,在BEA-test数据集上,该方法实现了单模型系统的最新最佳结果。同时,本文对最常见的英语语法纠错数据集进行了去分词处理,以更贴近自然文本书写方式。在此过程中发现原数据集中存在标注错误。通过实验分析了在去分词数据集上训练的影响,并评估了使用经修正错误样本的数据集所带来的效果。为促进研究可复现性,本文已公开训练所用源代码。

原文摘要 · Abstract (English)

Decoder-only large language models have shown superior performance in the fluency-edit English Grammatical Error Correction, but their adaptation for minimal-edit English GEC is still underexplored. To improve their effectiveness in the minimal-edit approach, we explore the error rate adaptation topic and propose a novel training schedule method. Our experiments set a new state-of-the-art result for a single-model system on the BEA-test set. We also detokenize the most common English GEC datasets to match the natural way of writing text. During the process, we find that there are errors in them. Our experiments analyze whether training on detokenized datasets impacts the results and measure the impact of the usage of the datasets with corrected erroneous examples. To facilitate reproducibility, we have released the source code used to train our models.

语法纠错大模型数据修复可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。