arXiv:2505.18091cs.LGcs.AI2025-05NeurIPS被引 14

大模型训练时数据混合比例变化会引发知识获取的突变现象

Data Mixing Can Induce Phase Transitions in Knowledge Acquisition

  • 通过控制实验发现,模型大小和数据混合比达到临界值时,知识记忆出现突然跃迁
  • 小模型在低混合比下几乎不记知识,超过阈值后迅速大量记忆;大模型在临界尺寸时从少量到全量记忆跃变
  • 揭示了模型容量分配机制,提出可预测相变的理论框架,适合模型训练调参研究者

大型语言模型通常在混合数据上训练:大部分来自网络爬取,小部分来自高质量领域知识数据集。本文表明,在此类混合数据训练中,知识从密集数据集获取的行为并非始终遵循平滑缩放规律,而是随模型规模和混合比例出现相变现象。在合成传记数据与网络爬取数据的受控实验中,我们发现:(1)当模型规模达到临界值时,模型突然从仅记忆极少传记转变为记忆绝大部分;(2)在临界混合比例以下,即使充分训练也几乎不记忆,而超过该阈值后则快速记忆更多传记。我们归因于容量分配机制——具有有限容量的模型需像背包问题求解器一样最小化整体测试损失,其最优数据分配可能随模型大小或混合比例发生突变。我们在信息论框架下形式化该直觉,发现这些相变可预测,临界混合比例与模型规模呈幂律关系。研究结果表明,大模型的最优混合配方未必适用于小模型,反之亦然。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are typically trained on data mixtures: most data come from web scrapes, while a small portion is curated from high-quality sources with dense domain-specific knowledge. In this paper, we show that when training LLMs on such data mixtures, knowledge acquisition from knowledge-dense datasets, unlike training exclusively on knowledge-dense data (arXiv:2404.05405), does not always follow a smooth scaling law but can exhibit phase transitions with respect to the mixing ratio and model size. Through controlled experiments on a synthetic biography dataset mixed with web-scraped data, we demonstrate that: (1) as we increase the model size to a critical value, the model suddenly transitions from memorizing very few to most of the biographies; (2) below a critical mixing ratio, the model memorizes almost nothing even with extensive training, but beyond this threshold, it rapidly memorizes more biographies. We attribute these phase transitions to a capacity allocation phenomenon: a model with bounded capacity must act like a knapsack problem solver to minimize the overall test loss, and the optimal allocation across datasets can change discontinuously as the model size or mixing ratio varies. We formalize this intuition in an information-theoretic framework and reveal that these phase transitions are predictable, with the critical mixing ratio following a power-law relationship with the model size. Our findings highlight a concrete case where a good mixing recipe for large models may not be optimal for small models, and vice versa.

大模型训练知识获取相变现象数据混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。