arXiv:2507.09404cs.LG2025-07NeurIPS被引 50

用数学规律确定大模型训练数据配比,省去盲目试错。

Scaling Laws for Optimal Data Mixtures

  • 基于缩放定律,建立数据混合比例与模型性能的预测模型。
  • 在三种大规模模型上验证,可准确预测不同数据配比下的损失值。
  • 仅需少量小规模实验,就能推断大模型在新配比下的表现。

大型基础模型通常在多个领域数据上训练,数据混合比例(各领域占比)对模型性能至关重要。传统选择方法依赖试错,大规模预训练下不切实际。本文提出一种系统性方法,利用缩放定律确定任意目标领域的最优数据混合比例。该方法能准确预测模型规模为 $N$、训练 $D$ 个标记(tokens)且采用特定领域权重向量 $h$ 时的损失。我们在三个不同且大规模的场景中验证了这些缩放定律的普适性:大语言模型(LLM)、原生多模态模型(NMM)和大视觉模型(LVM)的预训练。进一步证明,这些缩放定律可外推至新数据混合比例和不同规模:其参数可通过少量小规模训练运行准确估计,并用于预测更大规模及未见过的领域权重下的性能。缩放定律使得在给定训练预算($N$,$D$)下,为任意目标领域推导出最优领域权重成为可能,提供了一种比昂贵试错更合理的替代方案。

原文摘要 · Abstract (English)

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting this mixture relies on trial and error, which becomes impractical for large-scale pretraining. We propose a systematic method to determine the optimal data mixture for any target domain using scaling laws. Our approach accurately predicts the loss of a model of size $N$ trained with $D$ tokens and a specific domain weight vector $h$. We validate the universality of these scaling laws by demonstrating their predictive power in three distinct and large-scale settings: large language model (LLM), native multimodal model (NMM), and large vision models (LVM) pretraining. We further show that these scaling laws can extrapolate to new data mixtures and across scales: their parameters can be accurately estimated using a few small-scale training runs, and used to estimate the performance at larger scales and unseen domain weights. The scaling laws allow to derive the optimal domain weights for any target domain under a given training budget ($N$,$D$), providing a principled alternative to costly trial-and-error methods.

模型训练数据混合缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。