arXiv:2604.23627cs.CLcs.LG2026-04被引 20

首个罗马尼亚语语法纠错数据集,用预训练提升低资源语言纠错效果

Neural Grammatical Error Correction for Romanian

  • 构建1万对罗马尼亚语纠错语料,适配德国版错误标注工具
  • 预训练+微调使纠错准确率提升至F0.5 53.76(基线44.38)
  • 基于词性标注生成合成数据,可推广至其他低资源语言

非英语语言的语法纠错(GEC)资源匮乏,现有拼写检查器多仅限于简单规则。本文首次为罗马尼亚语构建包含1万对句子的GEC语料库,并将德国版ERRANT评分工具适配至罗马尼亚语以分析语料与提取修正操作。实验对比多种神经模型及预训练策略,证明其在低资源场景下有效。基线为仅在语料上训练的小型Transformer模型(F0.5为44.38),最佳模型则通过在人工生成数据上预训练大型Transformer,再在真实语料上微调,实现F0.5达53.76。所提合成数据生成方法仅需词性标注器,易于扩展至任意语言。

原文摘要 · Abstract (English)

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a first GEC corpus for Romanian consisting of 10k pairs of sentences. In addition, the German version of ERRANT (ERRor ANnotation Toolkit) scorer was adapted for Romanian to analyze this corpus and extract edits needed for evaluation. Multiple neural models were experimented, together with pretraining strategies, which proved effective for GEC in low-resource settings. Our baseline consists of a small Transformer model trained only on the GEC dataset (F0.5 of 44.38), whereas the best performing model is produced by pretraining a larger Transformer model on artificially generated data, followed by finetuning on the actual corpus (F0.5 of 53.76). The proposed method for generating additional training examples is easily extensible and can be applied to any language, as it requires only a POS tagger

语法纠错低资源语言预训练罗马尼亚语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。