arXiv:2507.11867cs.CL2025-07中稿 · CLNLP 2025

双向融合语法纠错与可接受性判断,提升多语言语法处理效果

COLA-GEC: A Bidirectional Framework for Enhancing Grammatical Acceptability and Error Correction

  • 用纠错数据增强可接受性模型,跨语言性能提升
  • 通过动态损失函数引入可接受性信号,引导纠错更合语法
  • 适合多语言语法研究者及自然语言处理开发者参考

语法错误纠正(GEC)和语法可接受性判断(COLA)是自然语言处理的核心任务,共享基础语法规则却常独立发展。本文提出COLA-GEC,一种双向框架,通过相互知识迁移同时提升两项任务。首先,利用GEC数据集增强可接受性模型,在多个语言上显著提升性能;其次,通过动态损失函数将可接受性信号融入GEC训练,有效引导纠错结果趋向语法正确输出。该方法在多个多语言基准上达到当前最优表现。全面的错误分析揭示了标点纠错仍是主要挑战,为未来语法建模提供改进方向。

原文摘要 · Abstract (English)

Grammatical Error Correction (GEC) and grammatical acceptability judgment (COLA) are core tasks in natural language processing, sharing foundational grammatical knowledge yet typically evolving independently. This paper introduces COLA-GEC, a novel bidirectional framework that enhances both tasks through mutual knowledge transfer. First, we augment grammatical acceptability models using GEC datasets, significantly improving their performance across multiple languages. Second, we integrate grammatical acceptability signals into GEC model training via a dynamic loss function, effectively guiding corrections toward grammatically acceptable outputs. Our approach achieves state-of-the-art results on several multilingual benchmarks. Comprehensive error analysis highlights remaining challenges, particularly in punctuation error correction, providing insights for future improvements in grammatical modeling.

语法纠错可接受性判断多语言双向学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。