arXiv:2506.13234cs.LG2025-06ICML被引 8

微小初始化差异可导致训练轨迹彻底偏离,揭示神经网络训练的混沌特性。

The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions

  • 通过参数扰动实验,发现训练初期轨迹对初始条件极度敏感。
  • 参数距离与表示相似性随训练时间迅速收敛,早期扰动影响持久。
  • 研究结果对微调、模型合并和集成学习有重要实践指导意义。

神经网络训练对初始化和随机梯度下降带来的随机性具有内在敏感性。然而,这种敏感性在多大程度上会导致模型权重或学习到的函数产生有意义的差异仍不明确。本文表明,在训练初期的“混沌”阶段,即使极小的扰动也会使原本相同的训练轨迹可靠地发散,且该效应随训练时间迅速减弱。我们通过(i)参数间的 $L^2$ 距离、(ii)模型间插值时的损失屏障、(iii)排列对齐后的参数 $L^2$ 距离与屏障、(iv)中间激活表示的相似性,量化了这种发散行为;揭示了不同超参数或微调设置下扰动如何引导训练轨迹走向不同的损失极小值。研究为神经网络训练稳定性提供了新见解,对微调、模型合并及模型集成多样性具有实际意义。

原文摘要 · Abstract (English)

Neural network training is inherently sensitive to initialization and the randomness induced by stochastic gradient descent. However, it is unclear to what extent such effects lead to meaningfully different networks, either in terms of the models' weights or the underlying functions that were learned. In this work, we show that during the initial "chaotic" phase of training, even extremely small perturbations reliably causes otherwise identical training trajectories to diverge-an effect that diminishes rapidly over training time. We quantify this divergence through (i) $L^2$ distance between parameters, (ii) the loss barrier when interpolating between networks, (iii) $L^2$ and barrier between parameters after permutation alignment, and (iv) representational similarity between intermediate activations; revealing how perturbations across different hyperparameter or fine-tuning settings drive training trajectories toward distinct loss minima. Our findings provide insights into neural network training stability, with practical implications for fine-tuning, model merging, and diversity of model ensembles.

神经网络训练稳定性混沌现象模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。