小规模开源大模型也能实现顶尖多语言翻译,效果媲美谷歌和GPT-4。
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
- 用少于100亿参数的开源模型做多语言翻译,系统评估六款模型表现。
- 提出新训练策略,使GemmaX2-28在28种语言上超越现有最优模型。
- 适合关注低成本高效率多语言翻译落地的研究者与开发者。
大规模语言模型(LLMs)展现出持续提升的多语言能力,即使小规模开源模型也实现了快速性能进步。本文系统研究了参数量小于十亿的开源大模型在多语言机器翻译(MT)任务中的表现。我们在六个主流开源模型上进行了全面评估,发现Gemma2-9B已具备出色的多语言翻译能力。随后,在持续预训练阶段引入‘先并行、后单语’(PFMS)数据混合策略,进一步提升翻译性能,推出GemmaX2-28——一个90亿参数模型,在28种语言上达到顶尖多语言翻译水平。具体而言,GemmaX2-28在所有测试语言上均优于当前最佳模型(如TowerInstruct和XALMA),且在多数语言上表现可媲美Google Translate和GPT-4-turbo。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of open LLMs with less than ten billion parameters to handle multilingual machine translation (MT) tasks. We conduct comprehensive evaluations on six popular LLMs and find that models like Gemma2-9B exhibit impressive multilingual translation capabilities. We then introduce the Parallel-First Monolingual-Second (PFMS) data mixing strategy in the continual pretraining stage to further enhance the MT performance and present GemmaX2-28, a 9B model achieving top-tier multilingual translation performance across 28 languages. Specifically, GemmaX2-28 consistently outperforms the state-of-the-art (SOTA) models such as TowerInstruct and XALMA and achieves competitive performance with Google Translate and GPT-4-turbo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。