arXiv:2602.02400cs.LG2026-02被引 1

实验证明噪声数据会引发大模型预训练损失发散,且与模型规模和噪声类型有关。

An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence

  • 通过注入可控噪声,系统研究噪声对预训练的影响机制。
  • 在480M至5.2B参数模型上,噪声显著增加损失发散概率。
  • 发现噪声引发的发散模式不同于高学习率导致的异常,可有效区分。

大规模预训练数据集推动了大语言模型(LLMs)的成功,但这些网络规模语料库不可避免地包含大量噪声数据,源于不受监管的网络内容或数据生成的随机性。尽管预训练者常推测此类噪声会导致训练不稳定甚至损失发散,但该现象仍缺乏深入理解。本文通过向原本干净的数据集注入受控的合成均匀随机噪声,系统研究了噪声是否引发大模型预训练发散及其机制。我们分析了从480M到5.2B参数的多个模型规模下的训练动态,结果表明:噪声数据确实会引起训练损失发散,且发散概率强烈依赖于噪声类型、噪声量级及模型规模。进一步发现,噪声诱导的发散具有与高学习率引发的发散不同的激活模式,并提供了可区分这两种失败模式的诊断方法。本研究为噪声数据如何影响大模型预训练损失发散提供了大规模、可控的实证刻画。

原文摘要 · Abstract (English)

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale LLM pretraining and, in the worst cases, loss divergence, this phenomenon remains poorly understood.In this work, we present a systematic empirical study of whether noisy data causes LLM pretraining divergences and how it does so. By injecting controlled synthetic uniformly random noise into otherwise clean datasets, we analyze training dynamics across model sizes ranging from 480M to 5.2B parameters. We show that noisy data indeed induces training loss divergence, and that the probability of divergence depends strongly on the noise type, amount of noise, and model scale. We further find that noise-induced divergences exhibit activation patterns distinct from those caused by high learning rates, and we provide diagnostics that differentiate these two failure modes. Together, these results provide a large-scale, controlled characterization of how noisy data affects loss divergence in LLM pretraining.

大模型训练噪声数据损失发散

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。