让大模型通晓低资源语言,跨语言能力全面增强
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
- 用多语种数据持续预训练Qwen2模型,提升跨语言理解
- 在多个评测集上超越现有模型,低资源语言性能显著提升
- 适合关注多语言AI落地、尤其是小语种应用的研究者
大型语言模型近年取得显著进展,但其优秀表现仍主要局限于英语等主流语言。许多模型在多语言任务中,特别是低资源语言上仍面临挑战。为此,我们提出了Marco-LLM:通过大规模多语言训练实现跨语言增强的通用语言模型。我们为多种低资源语言收集了大量多语言数据,并基于Qwen2模型进行了广泛的持续预训练,构建出名为Marco-LLM的多语言大模型。在包括MMMLU、AGIEval、Belebele、Flores-200、XCOPA在内的多个多语言基准上的综合评估显示,Marco-LLM在性能上显著优于当前最优模型。此外,该模型在任意语言间的机器翻译任务中也表现出明显提升,验证了其多语言能力的有效性。Marco-LLM不仅在低资源语言任务中表现优异,同时保持对英语等主流语言的强大性能,缩小了高资源与低资源语言间的能力差距。通过语言桥梁建设,体现了推动大模型在全球语言中准确运行的承诺。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challenges with multilingual tasks, especially when it comes to low-resource languages. To address this issue, we introduced Marco-LLM: Massive multilingual training for cross-lingual enhancement LLM. We have collected a substantial amount of multilingual data for several low-resource languages and conducted extensive continual pre-training using the Qwen2 models. This effort has resulted in a multilingual LLM named Marco-LLM. Through comprehensive evaluations on various multilingual benchmarks, including MMMLU, AGIEval, Belebele, Flores-200, XCOPA and many others, Marco-LLM has demonstrated substantial improvements over state-of-the-art LLMs. Furthermore, Marco-LLM achieved substantial enhancements in any-to-any machine translation tasks, showing the effectiveness of our multilingual LLM. Marco-LLM is a pioneering multilingual LLM designed to not only perform exceptionally well in multilingual tasks, including low-resource languages, but also maintain strong performance in English and other major languages, closing the performance gap between high- and low-resource language capabilities. By bridging languages, this effort demonstrates our dedication to ensuring LLMs work accurately across various languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。