研究随机权重下深度网络的固定点数量与稳定性,发现轻尾分布仅有一个稳定固定点。
Random weights of DNNs and emergence of fixed points
- 对比轻尾(如高斯)与重尾(如柯西)分布初始化的固定点数量
- 轻尾下几乎总存在唯一稳定固定点,重尾下出现多个稳定吸引子
- 固定点数量随网络深度呈现非单调变化,适合研究训练动力学者关注
本文研究输入输出维度相同的深层神经网络(DNNs),这类网络广泛应用于自编码器等场景。其训练过程可由固定点(FPs)表征。我们考察随机初始化权重矩阵分布对固定点数量及其稳定性的影响,重点关注独立同分布的轻尾(如高斯)与重尾(如柯西)分布。通过大量模拟发现:在轻尾分布下,多数架构均仅存在一个稳定固定点;而在重尾分布下,会涌现出多个固定点。这些固定点为稳定吸引子,其吸引域划分了输入空间。此外,固定点数量Q(L)随网络深度L呈现非单调变化。上述现象首先在未训练网络中观察到,并在训练过程中自然产生重尾分布后得到验证。
原文摘要 · Abstract (English)
This paper is concerned with a special class of deep neural networks (DNNs) where the input and the output vectors have the same dimension. Such DNNs are widely used in applications, e.g., autoencoders. The training of such networks can be characterized by their fixed points (FPs). We are concerned with the dependence of the FPs number and their stability on the distribution of randomly initialized DNNs' weight matrices. Specifically, we consider the i.i.d. random weights with heavy and light-tail distributions. Our objectives are twofold. First, the dependence of FPs number and stability of FPs on the type of the distribution tail. Second, the dependence of the number of FPs on the DNNs' architecture. We perform extensive simulations and show that for light tails (e.g., Gaussian), which are typically used for initialization, a single stable FP exists for broad types of architectures. In contrast, for heavy tail distributions (e.g., Cauchy), which typically appear in trained DNNs, a number of FPs emerge. We further observe that these FPs are stable attractors and their basins of attraction partition the domain of input vectors. Finally, we observe an intriguing non-monotone dependence of the number of fixed points $Q(L)$ on the DNNs' depth $L$. The above results were first obtained for untrained DNNs with two types of distributions at initialization and then verified by considering DNNs in which the heavy tail distributions arise in training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。