华为用迁移学习提升低资源印度语翻译质量,效果显著。
Machine Translation Advancements of Low-Resource Indian Languages by Transfer Learning
- 针对不同语言特点,分别采用微调和多语言建模策略。
- 英-阿萨姆语、英-曼尼普尔语最高达47.9 BLEU。
- 为低资源印度语翻译提供可复现的技术方案。
本文介绍了华为翻译中心(HW-TSC)在WMT24印度语言机器翻译共享任务中的提交方案。为构建低资源印度语言的可靠翻译系统,我们根据语言文字特征及现有开源模型支持情况,采用了两种不同的知识迁移策略。对于阿萨姆语(as)和曼尼普尔语(mn),我们在现有IndicTrans2开源模型基础上进行微调,实现英-这些语言间的双向翻译。对于卡西语(kh)和米佐语(mz),我们使用这四组双语数据及约8千条英-孟加拉语双语数据训练一个多语言基线模型,并进一步微调以实现英-卡西语、英-米佐语的双向翻译。迁移学习实验取得显著成果:en-as达23.5 BLEU,en-mn达31.8 BLEU,as-en达36.2 BLEU,mn-en达47.9 BLEU;多语言模型实验同样表现优异,en-kh达19.7 BLEU,en-mz达32.8 BLEU,kh-en达16.1 BLEU,mz-en达33.9 BLEU。这些结果不仅验证了迁移学习在低资源语言上的有效性,也推动了低资源印度语言机器翻译能力的发展。
原文摘要 · Abstract (English)
This paper introduces the submission by Huawei Translation Center (HW-TSC) to the WMT24 Indian Languages Machine Translation (MT) Shared Task. To develop a reliable machine translation system for low-resource Indian languages, we employed two distinct knowledge transfer strategies, taking into account the characteristics of the language scripts and the support available from existing open-source models for Indian languages. For Assamese(as) and Manipuri(mn), we fine-tuned the existing IndicTrans2 open-source model to enable bidirectional translation between English and these languages. For Khasi (kh) and Mizo (mz), We trained a multilingual model as a baseline using bilingual data from these four language pairs, along with an additional about 8kw English-Bengali bilingual data, all of which share certain linguistic features. This was followed by fine-tuning to achieve bidirectional translation between English and Khasi, as well as English and Mizo. Our transfer learning experiments produced impressive results: 23.5 BLEU for en-as, 31.8 BLEU for en-mn, 36.2 BLEU for as-en, and 47.9 BLEU for mn-en on their respective test sets. Similarly, the multilingual model transfer learning experiments yielded impressive outcomes, achieving 19.7 BLEU for en-kh, 32.8 BLEU for en-mz, 16.1 BLEU for kh-en, and 33.9 BLEU for mz-en on their respective test sets. These results not only highlight the effectiveness of transfer learning techniques for low-resource languages but also contribute to advancing machine translation capabilities for low-resource Indian languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。