用开源大模型实现46种语言的高质量机器翻译,性能媲美谷歌等商业系统。
Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models
- 基于Gemma3模型,通过持续预训练和指令微调提升多语言翻译能力。
- MiLMMT-46在46种语言上表现顶尖,超越多个SOTA模型。
- 适合关注低成本高性价比多语言翻译方案的研究者与开发者。
近年来,开源大语言模型在多语言能力方面持续进步。本文研究了开源LLM在多语言机器翻译(MT)中的应用,探讨了模型规模和数据规模对通过持续预训练与指令微调适配多语言MT的影响。基于Gemma3模型家族,我们构建了MiLMMT-46,在46种语言上实现了顶级的多语言翻译性能。大量实验表明,MiLMMT-46持续优于近期SOTA模型(如Seed-X、HY-MT-1.5、TranslateGemma),并达到与谷歌翻译、Gemini 3 Pro等强大专有系统相竞争的水平。模型已发布于https://huggingface.co/collections/xiaomi-research/milmmt-46,代码开源于https://github.com/xiaomi-research/gemmax。
原文摘要 · Abstract (English)
Open large language models (LLMs) have demonstrated improving multilingual capabilities in recent years. In this paper, we present a study of open LLMs for multilingual machine translation (MT) across a range of languages, and investigate the effects of model scaling and data scaling when adapting open LLMs to multilingual MT through continual pretraining and instruction finetuning. Based on the Gemma3 model family, we develop MiLMMT-46, which achieves top-tier multilingual translation performance across 46 languages. Extensive experiments show that MiLMMT-46 consistently outperforms recent state-of-the-art (SOTA) models, including Seed-X, HY-MT-1.5, and TranslateGemma, and achieves competitive performance with strong proprietary systems such as Google Translate and Gemini 3 Pro. Models are released at https://huggingface.co/collections/xiaomi-research/milmmt-46. Codes are released at https://github.com/xiaomi-research/gemmax.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。