arXiv:2509.14504cs.CLcs.AI2025-09被引 3

发布多语言语法纠错数据集OmniGEC,覆盖11种语言,助力跨语言语法修正研究。

Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction

  • 基于维基编辑与社交媒体文本构建多语言银标准数据集
  • 用GPT-4o-mini自动修正数据,经人工与自动评估确保质量
  • 在11种语言上微调模型达当前最佳性能,适合多语言NLP研究者

本文介绍OmniGEC,一个涵盖11种语言(捷克语、英语、爱沙尼亚语、德语、希腊语、冰岛语、意大利语、拉脱维亚语、斯洛文尼亚语、瑞典语、乌克兰语)的多语言语法错误修正银标准数据集。数据源自三个来源:目标语言的维基百科编辑、对应语言的Reddit子版块以及仅含乌克兰语的UberText 2.0社交文本语料库。其中维基百科编辑采用人工修正,而Reddit与UberText 2.0数据则通过GPT-4o-mini模型自动纠正。数据集的校正质量经自动与人工双重评估。最后,我们在OmniGEC数据集上微调两款开源大模型——Aya-Expanse(8B)和Gemma-3(12B),在段落级多语言语法纠错任务上取得当前最优结果。数据集与最优模型已公开于Hugging Face。

原文摘要 · Abstract (English)

In this paper, we introduce OmniGEC, a collection of multilingual silver-standard datasets for the task of Grammatical Error Correction (GEC), covering eleven languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Slovene, Swedish, and Ukrainian. These datasets facilitate the development of multilingual GEC solutions and help bridge the data gap in adapting English GEC solutions to multilingual GEC. The texts in the datasets originate from three sources: Wikipedia edits for the eleven target languages, subreddits from Reddit in the eleven target languages, and the Ukrainian-only UberText 2.0 social media corpus. While Wikipedia edits were derived from human-made corrections, the Reddit and UberText 2.0 data were automatically corrected with the GPT-4o-mini model. The quality of the corrections in the datasets was evaluated both automatically and manually. Finally, we fine-tune two open-source large language models - Aya-Expanse (8B) and Gemma-3 (12B) - on the multilingual OmniGEC corpora and achieve state-of-the-art (SOTA) results for paragraph-level multilingual GEC. The dataset collection and the best-performing models are available on Hugging Face.

语法纠错多语言数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。