arXiv:2409.11272cs.CLcs.AI2024-09被引 5

LOLA是支持160多种语言的开源大模型,高效处理多语言任务。

LOLA -- An Open-Source Massively Multilingual Large Language Model

  • 采用稀疏专家混合架构,提升多语言训练效率。
  • 在多语言生成与理解任务中表现优异,支持160+语言。
  • 开源可复现,适合研究多语言模型与资源受限场景。

本文介绍LOLA,一个基于稀疏专家混合Transformer架构、在超过160种语言上训练的大规模多语言语言模型。其架构与实现设计解决了利用语言多样性的同时保持效率并避免多语言常见问题的挑战。评估结果显示,该模型在自然语言生成与理解任务中表现具有竞争力。此外,我们展示了所学的专家路由机制能够利用隐含的语系演化模式,有望缓解多语言困境。本文深入分析了训练过程、数据集,并全面探讨了模型的优势与局限性。作为开源模型,LOLA促进可复现性,为未来研究提供坚实基础。研究结果有助于开发计算高效、跨语言表现强且可扩展的多语言模型。

原文摘要 · Abstract (English)

This paper presents LOLA, a massively multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. Our architectural and implementation choices address the challenge of harnessing linguistic diversity while maintaining efficiency and avoiding the common pitfalls of multilinguality. Our analysis of the evaluation results shows competitive performance in natural language generation and understanding tasks. Additionally, we demonstrate how the learned expert-routing mechanism exploits implicit phylogenetic linguistic patterns to potentially alleviate the curse of multilinguality. We provide an in-depth look at the training process, an analysis of the datasets, and a balanced exploration of the model's strengths and limitations. As an open-source model, LOLA promotes reproducibility and serves as a robust foundation for future research. Our findings enable the development of compute-efficient multilingual models with strong, scalable performance across languages.

多语言模型开源模型专家混合LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。