arXiv:2503.21023cs.LG2025-03NeurIPS被引 11

用概率方法优化大模型训练数据组合,省时省力还更准。

Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework

  • 基于多保真度贝叶斯优化,自动选择最优数据混合与训练参数
  • 在20M到10亿参数模型上实现2.6至3.3倍提速
  • 适合想高效训练大模型的研究者和工程团队

精心筛选数据源能显著提升大模型预训练效果,但现有方法多依赖直觉或高成本试错,难以跨领域推广。尽管规模定律提供了系统化思路,但传统确定性外推需强假设,其脆弱性已被前人指出。本文提出一种概率外推框架,避免刚性假设并显式建模性能不确定性。将数据配置优化建模为多保真度、多尺度贝叶斯优化的序贯决策问题,自适应选择数据混合、模型规模与训练步数,以平衡成本与信息增益。该框架可利用低成本实验中的噪声信息指导高成本训练决策。为加速进展,我们基于SlimPajama数据集构建了包含472次预训练运行的模拟器。结果显示,即使使用简单核函数与采集函数,也能在20M至10亿参数模型上实现2.6倍和3.3倍于多保真度贝叶斯优化与随机搜索的加速。结果表明,发展系统化、可迁移的数据混合优化方法具有显著效率潜力。

原文摘要 · Abstract (English)

Careful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, making them difficult to generalize across different data domains and downstream tasks. Although scaling laws can provide a principled and general approach for data curation, standard deterministic extrapolation from small-scale experiments to larger scales requires strong assumptions on the reliability of such extrapolation, whose brittleness has been highlighted in prior works. In this paper, we introduce a $\textit{probabilistic extrapolation framework}$ for data mixture optimization that avoids rigid assumptions and explicitly models the uncertainty in performance across decision variables. We formulate data curation as a sequential decision-making problem$\unicode{x2013}$multi-fidelity, multi-scale Bayesian optimization$\unicode{x2013}$where $\{$data mixtures, model scale, training steps$\}$ are adaptively selected to balance training cost and potential information gain. Our framework naturally gives rise to algorithm prototypes that leverage noisy information from inexpensive experiments to systematically inform costly training decisions. To accelerate methodological progress, we build a simulator based on 472 language model pre-training runs with varying data compositions from the SlimPajama dataset. We observe that even simple kernels and acquisition functions can enable principled decisions across training models from 20M to 1B parameters and achieve $\textbf{2.6x}$ and $\textbf{3.3x}$ speedups compared to multi-fidelity BO and random search baselines. Taken together, our framework underscores potential efficiency gains achievable by developing principled and transferable data mixture optimization methods.

大模型训练数据优化贝叶斯优化高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。