arXiv:2506.07719cs.CL2025-06被引 5

构建跨语言语法纠错标注框架,兼顾一致性与语言特性灵活性。

Multilingual Grammatical Error Annotation: Combining Language-Agnostic Framework with Language-Specific Flexibility

  • 采用无语言依赖基础+语言特化扩展的模块化设计
  • 支持英德捷韩中等多语言标注,覆盖通用与定制需求
  • 提升多语言语法纠错评估的一致性与可解释性

语法错误纠正(GEC)依赖准确的错误标注与评估,但现有框架如 errant 在拓展至语系多样语言时存在局限。本文提出一种标准化、模块化的多语言语法错误标注框架,结合语言无关基础与结构化语言特化扩展,实现跨语言的一致性与灵活性。通过使用 stanza 重实现 errant,扩大了多语言支持范围,并在英语、德语、捷克语、韩语和汉语上验证了该框架的适应性,涵盖从通用标注到精细语言学调整的应用。该工作推动了多语言 GEC 标注的可扩展性与可解释性,促进了多语言环境下的统一评估。完整代码与标注工具见 https://github.com/open-writing-evaluation/jp_errant_bea。

原文摘要 · Abstract (English)

Grammatical Error Correction (GEC) relies on accurate error annotation and evaluation, yet existing frameworks, such as $\texttt{errant}$, face limitations when extended to typologically diverse languages. In this paper, we introduce a standardized, modular framework for multilingual grammatical error annotation. Our approach combines a language-agnostic foundation with structured language-specific extensions, enabling both consistency and flexibility across languages. We reimplement $\texttt{errant}$ using $\texttt{stanza}$ to support broader multilingual coverage, and demonstrate the framework's adaptability through applications to English, German, Czech, Korean, and Chinese, ranging from general-purpose annotation to more customized linguistic refinements. This work supports scalable and interpretable GEC annotation across languages and promotes more consistent evaluation in multilingual settings. The complete codebase and annotation tools can be accessed at https://github.com/open-writing-evaluation/jp_errant_bea.

语法纠错多语言标注框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。