arXiv:2410.12883cs.CLcs.LG2024-10ACL被引 27

提出多语言模型缩放定律,用语系采样比预测性能并优化训练。

Scaling Laws for Multilingual Language Models

  • 以语系为单位分析性能,发现损失仅取决于自身采样比例。
  • 推导出数据量、模型大小与采样比的幂律关系,可精准预测性能。
  • 小模型得出的最优采样比可推广到大模型,节省训练资源。

我们为通用解码器型多语言语言模型提出一种新型缩放定律,解决多语言预训练中语言平衡难题。由于跨语言迁移,分析单个语言性能困难,因此我们将关注点从单个语言转向语言家族。提出并验证一个假设:每个语言家族的测试交叉熵损失仅由其自身采样比例决定,与其他语言无关。这一洞察简化了多语言缩放分析,使其可扩展至任意数量语言。基于此假设,我们推导出性能与数据集规模、模型规模及采样比例之间的幂律关系,可预测多种组合下的表现,并推导不同模型规模下的最优采样比例。为验证该缩放定律的有效性与准确性,我们在23种语言(涵盖5个语系)上训练超过100个模型进行大规模实验。结果表明,基于小模型(85M参数)推导的最优采样比例能有效推广至大模型(1.2B参数),实现高效的大规模多语言语言模型训练。

原文摘要 · Abstract (English)

We propose a novel scaling law for general-purpose decoder-only language models (LMs) trained on multilingual data, tackling the problem of balancing languages during multilingual pretraining. A primary challenge in studying multilingual scaling is the difficulty of analyzing individual language performance due to cross-lingual transfer. To address this, we shift the focus from individual languages to language families. We introduce and validate a hypothesis that the test cross-entropy loss for each language family is determined solely by its own sampling ratio, independent of other languages in the mixture. This insight simplifies the complexity of multilingual scaling and make the analysis scalable to an arbitrary number of languages. Building on this hypothesis, we derive a power-law relationship that links performance with dataset size, model size and sampling ratios. This relationship enables us to predict performance across various combinations of the above three quantities, and derive the optimal sampling ratios at different model scales. To demonstrate the effectiveness and accuracy of our proposed scaling law, we perform a large-scale empirical study, training more than 100 models on 23 languages spanning 5 language families. Our experiments show that the optimal sampling ratios derived from small models (85M parameters) generalize effectively to models that are several orders of magnitude larger (1.2B parameters), offering a resource-efficient approach for multilingual LM training at scale.

多语言缩放定律语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。