用多语言+单双语数据预训练,提升刚果语翻译质量
Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation
- 结合多语言与单双语数据进行预训练
- 在刚果语上实现显著翻译性能提升
- 适合低资源语言研究者参考
本研究探讨了针对低资源语言的机器翻译模型预训练策略有效性。尽管涉及阿菲利卡语、斯瓦希里语和祖鲁语等多种语言,但重点构建了基于瑞德与阿特埃克斯(2021)提出的高资源语言预训练方法的刚果语翻译模型。通过一系列实验,考察了多语言预训练及在预训练阶段融合单语和双语数据的方法。结果表明,采用多语言预训练并结合单双语数据可显著提升翻译质量。该研究为低资源机器翻译提供了有效预训练策略,有助于缩小高资源与低资源语言间的性能差距。研究成果推动面向边缘化群体的更包容、准确的自然语言处理模型发展。代码与数据集已公开,除部分因公共可用性变更无法获取的数据外。
原文摘要 · Abstract (English)
This research article examines the effectiveness of various pretraining strategies for developing machine translation models tailored to low-resource languages. Although this work considers several low-resource languages, including Afrikaans, Swahili, and Zulu, the translation model is specifically developed for Lingala, an under-resourced African language, building upon the pretraining approach introduced by Reid and Artetxe (2021), originally designed for high-resource languages. Through a series of comprehensive experiments, we explore different pretraining methodologies, including the integration of multiple languages and the use of both monolingual and parallel data during the pretraining phase. Our findings indicate that pretraining on multiple languages and leveraging both monolingual and parallel data significantly enhance translation quality. This study offers valuable insights into effective pretraining strategies for low-resource machine translation, helping to bridge the performance gap between high-resource and low-resource languages. The results contribute to the broader goal of developing more inclusive and accurate NLP models for marginalized communities and underrepresented populations. The code and datasets used in this study are publicly available to facilitate further research and ensure reproducibility, with the exception of certain data that may no longer be accessible due to changes in public availability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。