arXiv:2607.11052cs.LGcs.CL2026-07被引 4

研究不同数据域组合如何协同提升大模型性能

Domain-Aware Scaling Laws Uncover Data Synergy

论文配图:Domain-Aware Scaling Laws Uncover Data Synergy
图 1 · 摘自论文原文
  • 通过分析多种开源大模型的预训练数据混合,量化数据域间的协同效应
  • 发现某些数据组合能显著提升数学推理能力,而部分组合反而降低性能
  • 适用于希望优化数据配比以提升模型表现的研究者与工程师

机器学习进展常归因于模型规模和数据量的扩大,但数据构成同样关键。实证发现,来自不同领域的数据混合会产生非平凡的交互作用:例如加入代码数据可提升数学推理能力,而某些组合则引入干扰导致性能下降。我们统称此类现象为数据协同效应,即多领域贡献既可能超过各领域独立贡献之和,也可能低于该和。本文在语言模型预训练中形式化并量化了这种数据协同效应。基于具有多样化预训练混合的开源大模型的观测差异,我们估计了直接的领域-任务协同效应(一个领域对另一任务性能的贡献)以及二阶领域-领域协同效应(需多个领域共现才能产生的能力)。该框架相比无领域感知的缩放定律预测更准确,并获得稳定协同效应估计。通过在预测最优和预测最差的数据混合上训练模型,验证了协同效应估计能正确预测性能排名。

原文摘要 · Abstract (English)

Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings repeatedly show that combining datasets from different domains yields nontrivial interactions. For instance, adding code improves mathematical reasoning, while certain mixtures introduce interference that reduces model performance. We refer to these effects collectively as data synergy, where the contribution of multiple domains exceeds or falls short of the sum of their isolated contributions. In this work, we formalize and quantify data synergy in language model pretraining. Leveraging observational variation across open-weight LLMs with diverse pretraining mixtures, we estimate both direct domain-to-benchmark synergy (how one domain contributes to performance on another) and a second-order domain-domain synergy (capabilities that require co-occurrence of multiple domains). Our framework improves predictive accuracy over domain-agnostic scaling laws and recovers stable synergy estimates. We validate these estimates by training models on predicted optimal and predicted anti-optimal mixtures and confirm that our synergy estimates correctly predict performance rankings.

数据协同大模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。