用合成错误数据训练模型,自动修复豪萨语书写错误。
Automatic Correction of Writing Anomalies in Hausa Texts
- 用合成噪声构建40万组错-正句对,训练纠错模型。
- M2M100等模型在纠错任务上表现最优,提升下游任务性能。
- 开源数据集与方法,助力低资源语言NLP发展。
豪萨语文本常存在字符误替换、空格错误等书写异常,影响自然语言处理应用。本文提出通过微调基于Transformer的模型实现自动纠错。利用多个公开来源收集语料,生成超过40万对带噪声与纯净的豪萨语句子对,以模拟真实书写错误。同时,微调多种多语言及非洲语言模型,包括M2M100、AfriTeVA、NCAIR1/N-ATLaS、UBC-NLP/cheetah-base,以及BART和T5的变体。实验表明,尽管规模较小且预训练方式不同,M2M100仍达到最佳效果;纠错显著提升文本分类、机器翻译、问答系统及大模型提示任务的表现。本研究提供方法、公开数据集及模型对比,推动豪萨语NLP能力提升,并为其他低资源语言提供可迁移经验。
原文摘要 · Abstract (English)
Hausa texts are often characterized by writing anomalies, such as incorrect character substitutions and spacing errors, which sometimes hinder natural language processing (NLP) applications. This paper presents an approach to automatically correct anomalies by finetuning transformer-based models. Using a corpus gathered from several public sources, we create a large-scale parallel dataset of over 400,000 noisy-clean Hausa sentence pairs by introducing synthetically generated noise to mimic realistic writing errors. In addition, we finetune several multilingual and African language models, including M2M100, AfriTeVA, NCAIR1/N-ATLaS, UBC-NLP/cheetah-base, and other variants of BART and T5 for this correction task. Our experimental results demonstrate that models such as M2M100 achieve state-of-the-art results despite their smaller size and distinct pretraining, and that correcting errors can have a significant impact in improving downstream tasks such as text classification, machine translation, question answering, and LLM prompting in general. This research provides a methodology, a publicly available dataset, and a comparison of models to improve Hausa text quality, thereby advancing NLP capabilities for the language and offering transferable insights for other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。