arXiv:2409.17892cs.CL2024-09被引 18

EMMA-500通过546种语言持续预训练,提升低资源语言表现。

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models

  • 用546种语言的MaLA语料持续预训练Llama 2 7B模型
  • 在多语言任务上显著提升跨语言迁移与适应能力
  • 适合需要多语言支持的研究者和开发者

本文提出EMMA-500,一个在546种语言文本上持续预训练的大规模多语言语言模型,旨在增强低资源语言的表现。为支持持续预训练,我们构建了MaLA语料库,整合了多个领域、经过筛选的多语言数据集。基于该语料库,我们对Llama 2 7B模型进行了大规模持续预训练,得到EMMA-500,其在涵盖多种多语言任务的广泛基准测试中表现出色。结果表明,持续预训练能有效扩展大语言模型的语言能力,尤其在代表性不足的语言上实现显著提升,推动跨语言迁移、任务泛化和语言适应性的进步。我们公开发布MaLA语料库、EMMA-500模型权重、脚本及生成结果。

原文摘要 · Abstract (English)

In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on texts across 546 languages designed for enhanced multilingual performance, focusing on improving language coverage for low-resource languages. To facilitate continual pre-training, we compile the MaLA corpus, a comprehensive multilingual dataset enriched with curated datasets across diverse domains. Leveraging this corpus, we conduct extensive continual pre-training of the Llama 2 7B model, resulting in EMMA-500, which demonstrates robust performance across a wide collection of benchmarks, including a comprehensive set of multilingual tasks. Our results highlight the effectiveness of continual pre-training in expanding large language models' language capacity, particularly for underrepresented languages, demonstrating significant gains in cross-lingual transfer, task generalization, and language adaptability. We release the MaLA corpus, EMMA-500 model weights, scripts, and model generations.

多语言模型持续预训练低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。