用实验设计方法优化大模型训练数据混合比例,提升效率与效果。
Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

- 将数据混合视为混合实验,用响应曲面建模各领域贡献。
- 通过稀疏Scheffé模型发现领域间交互作用显著提升效果。
- 设计更高效的代理实验方案,减少约25%测试次数仍可准确排序。
数据混合是大语言模型预训练中的核心设计问题:在固定词元预算下,需决定每个领域的数据分配比例。现有基于代理的方法通过在候选混合上训练小型模型、拟合响应模型,并利用其选择大规模训练的混合方案。本文揭示该流程具有经典混合实验的结构:数据领域为组分,词元占比为组分比例,代理训练运行为实验设计点,验证损失构成概率单纯形上的响应曲面。我们采用稀疏二阶Scheffé响应曲面模型,构建针对代理数据混合实验的模型稳健$$\mathcal{I}$-最优设计。以RegMix为实证案例,证明该框架既能解释观测到的混合响应,又能设计更高效的代理实验。Scheffé分析显示,领域价值具有强关联性:多个在加性效应下表现弱的领域,因与网络文本的成对交互而变得有利。稀疏Scheffé模型在不同模型规模下保持混合排名一致,且性能优于灵活的机器学习预测器,同时提供加性与交互效应的显式分解。在基于观测代理训练响应校准的模拟研究中,模型稳健$$\mathcal{I}$-最优设计在剔除约25%原始代理运行后,仍能恢复正确的混合排序。结果表明,大模型数据混合应被视作不仅是一个预测问题,更是一个实验设计问题,其中代理混合本身可被优化以提升统计效率。
原文摘要 · Abstract (English)
Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on candidate mixtures, fitting a response model, and using the response to select mixtures for larger-scale training. We show that this workflow has the structure of a classical mixture experiment. Under this view, data domains are mixture components, token shares are component proportions, proxy-training runs are experimental design points, and validation loss defines a response surface over the probability simplex. We develop this formulation using sparse second-order Scheffé response-surface models and construct model-robust $\mathcal{I}$-optimal designs for proxy data-mixing experiments. Using RegMix as an empirical case study, we demonstrate how the framework can both interpret observed mixture responses and design more efficient proxy experiments. The Scheffé analysis shows that domain value is strongly relational: several domains that are weak under additive effects become favourable through pairwise interactions, especially through combinations with web-derived text. The sparse Scheffé model preserves mixture rankings across model scales and remains competitive with a flexible machine-learning predictor while providing an explicit decomposition of additive and interaction effects. In a simulation study calibrated to observed proxy-training responses, model-robust $\mathcal{I}$-optimal designs recover the relevant mixture ordering after removing about 25\% of the original proxy runs. These results suggest that LLM data mixing should be treated not only as a prediction problem, but also as an experimental-design problem in which the proxy mixtures themselves can be chosen to improve statistical efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。