arXiv:2511.21700cs.CL2025-11AAAI

JELV自动评估语法纠错修正项有效性,提升评估准确性和模型泛化能力。

JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error Correction

  • 基于语法规则、忠实度和流畅性三维度,自动验证纠错修改的合理性
  • 在PEVData数据集上,大模型判别准确率达90%,轻量模型精度达85%
  • 可扩展单参考数据集,助力纠错模型性能提升,适合评估与训练场景

现有语法纠错(GEC)系统因参考答案多样性不足,导致评估结果偏低且模型泛化受限。为此,本文提出编辑级有效性评判器(JELV),从语法规则、忠实度和流畅性三个维度自动验证纠错修改。基于自建的人工标注成对编辑有效性数据集PEVData,JELV提供两种实现:多轮大模型判官流程与人类标注者达成90%一致,轻量化DeBERTa分类器在有效编辑上达到85%精确率。进一步,利用JELV重新判定评估中的误报,结合去耦误报与流畅性评分,构建综合评估指标,实现与人工判断的最先进相关性。同时,用JELV筛选大模型生成的纠错候选,扩展包含38,692个源句的BEA19单参考数据集。在此扩展数据集上重训主流GEC模型,取得可测量的性能提升。JELV为增强参考多样性、改进评估与模型泛化提供了可扩展方案。

原文摘要 · Abstract (English)

Existing Grammatical Error Correction (GEC) systems suffer from limited reference diversity, leading to underestimated evaluation and restricted model generalization. To address this issue, we introduce the Judge of Edit-Level Validity (JELV), an automated framework to validate correction edits from grammaticality, faithfulness, and fluency. Using our proposed human-annotated Pair-wise Edit-level Validity Dataset (PEVData) as benchmark, JELV offers two implementations: a multi-turn LLM-as-Judges pipeline achieving 90% agreement with human annotators, and a distilled DeBERTa classifier with 85% precision on valid edits. We then apply JELV to reclassify misjudged false positives in evaluation and derive a comprehensive evaluation metric by integrating false positive decoupling and fluency scoring, resulting in state-of-the-art correlation with human judgments. We also apply JELV to filter LLM-generated correction candidates, expanding the BEA19's single-reference dataset containing 38,692 source sentences. Retraining top GEC systems on this expanded dataset yields measurable performance gains. JELV provides a scalable solution for enhancing reference diversity and strengthening both evaluation and model generalization.

语法纠错评估方法数据扩展大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。