arXiv:2604.16380cs.CLcs.LG2026-04综述被引 2

优化大模型预训练数据配比,提升效率与泛化能力

Data Mixing for Large Language Models Pretraining: A Survey and Outlook

  • 将数据混合建模为概率单纯形上的双层优化问题
  • 提出静态与动态混合的细粒度分类体系
  • 适合关注大模型数据策略的研究者和工程师

大语言模型依赖于海量异构语料进行预训练,数据组成对训练效率和下游泛化能力有决定性影响。与样本级数据选择不同,数据混合通过优化领域级采样权重,在有限算力与数据预算下更有效地分配资源。近年来,一系列系统性数据混合方法被提出,但文献分散且缺乏专门综述。本文首次系统梳理大模型预训练中的数据混合方法:将数据混合优化形式化为概率单纯形上的双层问题,阐明其在预训练流程中的作用,并说明现有方法如何使该框架可计算。构建细粒度分类体系,按静态/动态混合划分,其中静态进一步分为规则与学习驱动,动态分为自适应与外部引导。分析各类代表性方法的性能-成本权衡。指出跨领域迁移性差、目标函数不统一、评估标准缺失等共性挑战。最后提出细粒度领域划分、反向数据混合及流水线感知设计等前瞻性方向,为未来研究提供思路。

原文摘要 · Abstract (English)

Large language models (LLMs) rely on pretraining on massive and heterogeneous corpora, where training data composition has a decisive impact on training efficiency and downstream generalization under realistic compute and data budget constraints. Unlike sample-level data selection, data mixing optimizes domain-level sampling weights to allocate limited budgets more effectively. In recent years, a growing body of work has proposed principled data mixing methods for LLM pretraining; however, the literature remains fragmented and lacks a dedicated, systematic survey. This paper provides a comprehensive review of data mixing for LLM pretraining. We first formalize data mixture optimization as a bilevel problem on the probability simplex and clarify the role of data mixing in the pretraining pipeline, and briefly explain how existing methods make this formulation tractable in practice. We then introduce a fine-grained taxonomy that organizes existing methods along two dimensions: static versus dynamic mixing. Static mixing is further categorized into rule-based and learning-based methods, while dynamic mixing is grouped into adaptive and externally guided families. For each class, we summarize representative approaches and analyze their strengths and limitations from a performance-cost trade-off perspective. Building on this analysis, we highlight challenges that cut across methods, including limited transferability across data domains, optimization objectives, models, and validation sets, as well as unstandardized evaluation protocols and benchmarks, and the inherent tension between performance gains and cost control in learning-based methods. Finally, we outline several exploratory directions, including finer-grained domain partitioning and inverse data mixing, as well as pipeline-aware designs, aiming to provide conceptual and methodological insights for future research.

大模型数据混合预训练综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。