提出新方法提升大模型训练中的数据混合效果
Rethinking Data Mixing from the Perspective of Large Language Models
- 从梯度动态角度建立领域与训练关系的理论框架
- 在不同规模GPT-2上验证,性能持续优于基线
- 适合关注大模型训练优化的研究者和工程师
数据混合策略对大语言模型训练至关重要。实证表明,不恰当的策略会显著降低模型泛化能力。尽管近期方法提升了经验性能,但若干根本问题仍待解答:何为一个领域?人类与模型对领域的感知是否一致?领域加权如何影响泛化?本文通过建立梯度动态与领域分布之间的形式化联系,提出理论框架,阐明领域在训练中的作用。基于此分析,我们提出DoGraph,一种将数据调度建模为图约束优化问题的重加权框架。在多个尺度的GPT-2模型上进行的大量实验表明,DoGraph始终表现出竞争力。
原文摘要 · Abstract (English)
Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly reduce generalization. Although recent methods have improved empirical performance, several fundamental questions remain open: what constitutes a domain, whether human and model perceptions of domains are aligned, and how domain weighting influences generalization. We address these questions by establishing formal connections between gradient dynamics and domain distributions, offering a theoretical framework that clarifies the role of domains in training dynamics. Building on this analysis, we introduce DoGraph, a reweighting framework that formulates data scheduling as a graph-constrained optimization problem. Extensive experiments on GPT-2 models of varying scales demonstrate that DoGraph consistently achieves competitive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。