arXiv:2607.27928cs.LG2026-07

用贝叶斯方法优化多领域数据权重,更稳更快地提升大模型性能。

Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

论文配图:Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
图 1 · 摘自论文原文
  • 引入伽马先验的狄利克雷分布,从观测中推断数据权重
  • 在少于搜索法1/10数据下实现稳定高效权重学习
  • 适合大规模训练中需要自动调优数据混合比的研究者

大语言模型的性能受多领域预训练数据分布的影响。早期依赖人工经验,但随着数据复杂度上升,难以捕捉领域间的协同效应。现有主流方法通过拟合域权重与验证损失的映射函数来寻找最优权重,但依赖秩不变性或缩放律等强假设,常导致显著估计偏差。直接从数据优化权重方案虽具潜力,却存在优化轨迹不稳定、计算开销过大等问题。本文提出一种贝叶斯域加权方法,通过引入从观测中学习的伽马先验信息,在狄利克雷分布框架下推断域权重。实验表明,该方法可实现稳定高效的权重学习,在远少于基于函数拟合的搜索方法所需数据量下,识别出最优数据混合比例,为大规模应用重燃了基于优化的域权重设计可能性。

原文摘要 · Abstract (English)

The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation losses, and then find the optimal domain weights to minimize validation losses. These methods rely on strong structural assumptions, such as rank invariance or scaling laws, which are often violated, resulting in non-negligible estimation bias. A promising approach is to directly optimize the weighting scheme from data. However, it suffers from unstable optimization trajectory and prohibitive computational overhead, limiting its potential to search better domain weights configurations. This paper presents a Bayesian domain weighting method to infer the weights from a Dirichlet distribution via introducing Gamma prior information learned from observations. Experimental results demonstrate that proposed method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale applications.

贝叶斯优化数据混合大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。