用机器翻译模型提升西非语言扎尔马的语法纠错能力
Grammatical Error Correction for Low-Resource Languages: The Case of Zarma
- 采用M2M100机器翻译模型进行语法纠错
- 检测率95.82%,建议准确率78.90%,人工评分3.0/5.0
- 适合低资源语言语法纠错研究者参考
语法错误纠正(GEC)旨在提升文本质量和可读性。以往研究主要聚焦高资源语言,而低资源语言缺乏有效工具。为此,我们针对西非使用人数超过五百万的扎尔马语开展GEC研究,比较了规则方法、机器翻译(MT)模型和大语言模型(LLMs)三种方案。基于包含超过25万条样本的合成与人工标注数据集评估,结果表明基于M2M100的MT方法表现最优,在自动评估中检测率达95.82%,建议准确率为78.90%,人工评估(来自母语者)平均得分为3.0/5.0(语法与逻辑修正)。规则方法仅对拼写错误有效,无法处理复杂上下文错误;Gemma 2b与MT5-small表现中等。该结论在另一西非语言巴姆巴拉上得到验证,支持采用MT模型推动低资源语言的GEC发展。
原文摘要 · Abstract (English)
Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource languages lack robust tools. To address this shortcoming, we present a study on GEC for Zarma, a language spoken by over five million people in West Africa. We compare three approaches: rule-based methods, machine translation (MT) models, and large language models (LLMs). We evaluated GEC models using a dataset of more than 250,000 examples, including synthetic and human-annotated data. Our results showed that the MT-based approach using M2M100 outperforms others, with a detection rate of 95.82% and a suggestion accuracy of 78.90% in automatic evaluations (AE) and an average score of 3.0 out of 5.0 in manual evaluation (ME) from native speakers for grammar and logical corrections. The rule-based method was effective for spelling errors but failed on complex context-level errors. LLMs -- Gemma 2b and MT5-small -- showed moderate performance. Our work supports use of MT models to enhance GEC in low-resource settings, and we validated these results with Bambara, another West African language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。