用标量嵌入法压缩训练轨迹,看清神经网络动态演化规律。
Scalar Representations of Neural Network Training Dynamics

- 将训练过程视为时间网络,通过标量嵌入降维表示。
- 保留了敏感性、李雅普诺夫指数等关键动态特征。
- 可定义特征时间,适合研究优化轨迹的统计行为。
人工神经网络的训练可视为高维损失空间中的轨迹演化,但参数量庞大导致直接分析困难。本文将训练轨迹建模为时序网络,采用近期提出的时序网络标量嵌入方法,探究其能否提供有意义的低维表征。以在MNIST分类任务上训练的多层感知机为例,结果表明该嵌入能有效保留原始参数空间中的主要动态特征,包括特定学习率区间下对初值敏感性的出现,以及网络最大李雅普诺夫指数的准确重建。进一步利用嵌入后的标量轨迹定义了一种类似李雅普诺夫时间的特征时间,标志着初始接近轨迹间指数分离趋于饱和的时刻,反映原高维系统中轨迹的典型去相关时间。最后,通过嵌入空间中的间距可观测量研究渐近训练状态的统计结构,发现经缩放后的渐近间距分布跨不同初值收敛于同一形式,与偏斜对数正态分布相容。总体表明,标量低维嵌入为研究和可视化神经网络优化轨迹的动力学性质提供了有效框架。
原文摘要 · Abstract (English)
Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging. In this work, we treat such training trajectories as temporal networks and apply recently proposed strategies for the scalar embedding of temporal networks. We investigate whether such a scalar embedding provides a meaningful low-dimensional representation of neural network training dynamics. Using a multilayer perceptron trained on the MNIST classification task, we show that the embedding preserves the main dynamical features observed in the original parameter space, including the emergence of sensitivity to initial conditions for specific learning rate regimes and an accurate reconstruction of the network's maximum Lyapunov exponent. We then use the embedded scalar trajectory to define a characteristic time, analogous to a Lyapunov time, after which the exponential separation between initially close embedded trajectories saturates. This characteristic time captures the typical decorrelation time between initially close network trajectories in the original high-dimensional system. Finally, we investigate the statistical organization of asymptotic training states through a spacing observable defined in the embedded space. We find that the distributions of rescaled asymptotic spacings collapse onto a common form across initial conditions and are compatible with a skew lognormal distribution. Altogether, our results suggest that scalar low-dimensional embeddings provide a useful framework for studying and visualizing the dynamical properties of neural network optimization trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。