arXiv:2602.15210cs.LG2026-02被引 2

通过精细化语言数据筛选,用少量高质量数据实现高效多语言模型训练。

ÜberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset

  • 针对每种语言独立优化数据质量,提升整体多语言性能。
  • 仅用总数据量8%的精选内容,训练出性能媲美大模型的多语言系统。
  • 适合追求低算力消耗、高多语言能力的模型开发者。

多语言是现代基础模型的核心能力,但不同语言数据分布不均,联合训练常引发性能干扰(即“多语言诅咒”)。我们研究了13种语言的数据筛选策略,发现多数性能下降并非由多语言扩展本身导致,而是可修复的数据质量问题。在受控双语实验中,任一语言的数据优化均能带动其他语言表现:优化英语使13种语言中有12种非英语语言性能提升;反之亦然。定制化每语言数据筛选带来显著的单语性能增长。将此方法扩展至大规模通用训练数据,我们发现仅占总量8%的精选多语言数据仍极为有效。我们据此构建了一个20万亿令牌的预训练语料库,全部来自公开数据。使用其中1万亿令牌随机子集训练的30亿和80亿参数模型,在多语言准确率上达到领先水平,且训练计算量仅为现有强基线的1/4至1/10,建立新的多语言性能与算力权衡前沿。该成果也应用于4000亿参数的Trinity Large(400B/A13B)模型,使其在同等训练算力下展现出优异的多语言表现。结果表明,针对性的逐语言数据筛选可缓解多语言干扰,实现高效多语言扩展。

原文摘要 · Abstract (English)

Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across languages. A further challenge is the performance interference that can arise from joint multilingual training, commonly referred to as the "curse of multilinguality". We study multilingual data curation across thirteen languages and find that many reported regressions are not inherent to multilingual scaling but instead stem from correctable deficiencies in data quality and composition rather than fundamental capacity limits. In controlled bilingual experiments, improving data quality for any single language benefits others: curating English improves non-English performance in 12 of 13 languages, while curating non-English yields reciprocal improvements in English. Bespoke per-language curation produces substantially larger within-language improvements. Extending these findings to large-scale general-purpose training mixtures, we show that curated multilingual allocations comprising under 8% of total tokens remain remarkably effective. We operationalize this approach within an effort that produced a 20T-token pretraining corpus derived entirely from public sources. Models with 3B and 8B parameters trained on a 1T-token random subset achieve competitive multilingual accuracy with 4-10x fewer training FLOPs than strong public baselines, establishing a new Pareto frontier in multilingual performance versus compute. Moreover, these benefits extend to frontier model scale: the 20T-token corpus served as part of the pretraining dataset for Trinity Large (400B/A13B), which exhibits strong multilingual performance relative to its training FLOPs. These results show that targeted, per-language data curation mitigates multilingual interference and enables compute-efficient multilingual scaling.

多语言数据筛选高效训练预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。