利用跨语言知识迁移,提升斯拉夫语系低资源语言翻译效果
MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality
- 通过跨语言知识迁移增强多语言翻译模型性能
- 在零样本场景下仍实现低资源斯拉夫语高效翻译
- 开源了最先进的斯拉夫语NMT模型,适合语言研究者使用
多语言神经机器翻译是否导致多语言诅咒,还是能在语言家族内实现跨语言知识迁移?本研究探索了扩展NMT数据范围的多种方法,并证明即使在零样本翻译条件下,低资源语言也能获得跨语言收益。本文提供了针对部分斯拉夫语之间的最先进开源NMT模型。这些模型已发布于HuggingFace Hub(https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683),采用CC BY 4.0许可。斯拉夫语族包含形态丰富的中欧与东欧语言,尽管母语使用者达数亿,但当前斯拉夫语神经机器翻译研究仍显不足。近期多数NMT研究集中于英语、西班牙语、德语等高资源语言(如WMT23通用翻译任务中8个方向中有7个涉及英语),或覆盖多个语族的大规模多语言模型,以及评估技术。
原文摘要 · Abstract (English)
Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for extending the available data-regime in NMT and we prove cross-lingual benefits even in 0-shot translation regime for low-resource languages. With this paper, we provide state-of-the-art open-source NMT models for translating between selected Slavic languages. We released our models on the HuggingFace Hub (https://hf.co/collections/allegro/multislav-6793d6b6419e5963e759a683) under the CC BY 4.0 license. Slavic language family comprises morphologically rich Central and Eastern European languages. Although counting hundreds of millions of native speakers, Slavic Neural Machine Translation is under-studied in our opinion. Recently, most NMT research focuses either on: high-resource languages like English, Spanish, and German - in WMT23 General Translation Task 7 out of 8 task directions are from or to English; massively multilingual models covering multiple language groups; or evaluation techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。