用多语言对齐数据训练大模型,显著提升低资源语言表现。
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
- 构建113语言的多向对齐语料TED2025,最多支持50语言并行。
- 在6个基准上,对齐数据训练模型性能全面优于非对齐数据。
- 适合想提升多语言模型跨语言理解能力的研究者和开发者。
持续预训练和指令微调大规模多语言数据已被证明能有效扩展大语言模型(LLM)在低资源语言上的能力。然而,此类数据的非对齐特性限制了其捕捉跨语言语义的能力。相比之下,多向平行数据(相同内容在多种语言中对齐)具有更强的跨语言一致性,更有利于提升多语言性能。本文基于TED演讲构建了一个大规模、高质量的多向平行语料库TED2025,覆盖113种语言,最多支持50种语言并行对齐,确保广泛的语言覆盖。利用该数据集,我们研究了优化多向平行数据使用的方法,包括持续预训练、指令微调策略及关键影响因素分析。在六个多语言基准上的实验表明,基于多向平行数据训练的模型始终优于基于非对齐多语言数据训练的模型。
原文摘要 · Abstract (English)
Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance. In this paper, we introduce a large-scale, high-quality multi-way parallel corpus, TED2025, based on TED Talks. The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage. Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors. Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。