arXiv:2506.00469cs.CL2025-06被引 6

用双语翻译数据让大模型掌握500种语言,尤其提升低资源语言表现。

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

  • 用2500+语对构建双语语料库,持续预训练大模型。
  • 在7项任务12个基准上,低资源语言性能显著提升。
  • 适合多语言研究者和需要跨语言能力的开发者。

本文研究大规模多语言持续预训练中是否使用平行语料的关键设计问题。针对Llama3系列模型扩展至500种语言的多语言适应,我们构建了包含超过2500个语言对的MaLA双语翻译语料库。基于此,我们开发了EMMA-500 Llama 3系列四个大规模多语言模型——从基础Llama3模型持续预训练,覆盖多样数据组合,总计达6710亿词元。通过对比有无双语翻译数据的持续预训练效果,我们在7项任务和12个基准上全面评估发现,双语数据能有效提升语言迁移能力和整体性能,尤其在低资源语言上表现更优。相关语料库、模型、代码与生成结果均已开源。

原文摘要 · Abstract (English)

This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for massively multilingual language adaptation of the Llama3 family of models to 500 languages. To this end, we construct the MaLA bilingual translation corpus, containing data from more than 2,500 language pairs. Subsequently, we develop the EMMA-500 Llama 3 suite of four massively multilingual models -- continually pre-trained from the Llama 3 family of base models extensively on diverse data mixes up to 671B tokens -- and explore the effect of continual pre-training with or without bilingual translation data. Comprehensive evaluation across 7 tasks and 12 benchmarks demonstrates that bilingual data tends to enhance language transfer and performance, particularly for low-resource languages. We open-source the MaLA corpus, EMMA-500 Llama 3 suite artefacts, code, and model generations.

多语言大模型翻译数据持续预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。