提出ATLAS模型,解决多语言训练中的规模扩展难题。
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- 构建跨语言迁移矩阵,量化1444对语言间的互惠效益。
- 发现不牺牲性能的前提下扩展语言的最优模型与数据配比。
- 明确从头预训练或微调的计算拐点,适合多语言研发者参考。
现有规模定律研究主要聚焦英语,但主流AI模型服务于数十亿国际用户。本文开展迄今最大规模的多语言规模定律研究,共完成774次多语言训练实验,覆盖1000万至80亿参数、400多种训练语言及48种评估语言。提出自适应迁移规模定律(ATLAS),在单语和多语预训练中均显著优于现有方法,外推泛化能力提升超0.3 R²。通过分析揭示多语言学习动态、语言间迁移特性及多语言诅咒机制:首先构建跨语言迁移矩阵,实证测量38×38=1444个语言对间的互惠得分;其次提出无语言依赖的规模定律,阐明添加语言时优化模型大小与数据量的方法;最后识别出从头预训练与基于多语检查点微调的计算拐点。研究成果为跨语言规模化提供科学基础,助力超越以英语为中心的高效模型扩展。
原文摘要 · Abstract (English)
Scaling laws research has focused overwhelmingly on English -- yet the most prominent AI models explicitly serve billions of international users. In this work, we undertake the largest multilingual scaling laws study to date, totaling 774 multilingual training experiments, spanning 10M-8B model parameters, 400+ training languages and 48 evaluation languages. We introduce the Adaptive Transfer Scaling Law (ATLAS) for both monolingual and multilingual pretraining, which outperforms existing scaling laws' out-of-sample generalization often by more than 0.3 R^2. Our analyses of the experiments shed light on multilingual learning dynamics, transfer properties between languages, and the curse of multilinguality. First, we derive a cross-lingual transfer matrix, empirically measuring mutual benefit scores between 38 x 38=1444 language pairs. Second, we derive a language-agnostic scaling law that reveals how to optimally scale model size and data when adding languages without sacrificing performance. Third, we identify the computational crossover points for when to pretrain from scratch versus finetune from multilingual checkpoints. We hope these findings provide the scientific foundation for democratizing scaling laws across languages, and enable practitioners to efficiently scale models -- beyond English-first AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。