arXiv:2510.25947cs.CLcs.AI2025-10被引 10

多语言数据混合可提升模型能力,不降性能也不必牺牲低资源语言表现

Revisiting Multilingual Data Mixtures in Language Model Pretraining

  • 用均衡的多语言数据训练,避免性能下降
  • 英语作枢纽语言能跨语系提升效果,非同族语言选枢纽更优
  • 400种语言下无显著性能衰减,适合构建通用模型

大型语言模型预训练中的多语言数据混合策略长期存在争议,常担心语言覆盖广度与模型性能之间的权衡(即‘多语言诅咒’)。本文通过在25至400种语言的语料上训练11亿和30亿参数的语言模型,重新审视这些假设。研究发现:在每种语言有足够的训练文本量时,加入英文与其他语言的数据不会损害各语言的原生性能;以英语作为枢纽语言(高资源语言)可促进跨语系泛化,而从特定语族中选择枢纽语言并未显著提升该语族内语言的表现;在当前模型规模下,随着训练语言数量增加,并未观察到明显的‘多语言诅咒’现象。结果表明,合理平衡的多语言数据可增强模型能力,且不影响低资源语言表现。

原文摘要 · Abstract (English)

The impact of different multilingual data mixtures in pretraining large language models (LLMs) has been a topic of ongoing debate, often raising concerns about potential trade-offs between language coverage and model performance (i.e., the curse of multilinguality). In this work, we investigate these assumptions by training 1.1B and 3B parameter LLMs on diverse multilingual corpora, varying the number of languages from 25 to 400. Our study challenges common beliefs surrounding multilingual training. First, we find that combining English and multilingual data does not necessarily degrade the in-language performance of either group, provided that languages have a sufficient number of tokens included in the pretraining corpus. Second, we observe that using English as a pivot language (i.e., a high-resource language that serves as a catalyst for multilingual generalization) yields benefits across language families, and contrary to expectations, selecting a pivot language from within a specific family does not consistently improve performance for languages within that family. Lastly, we do not observe a significant "curse of multilinguality" as the number of training languages increases in models at this scale. Our findings suggest that multilingual data, when balanced appropriately, can enhance language model capabilities without compromising performance, even in low-resource settings

多语言预训练语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。