研究神经网络权重在训练中如何逼近高斯分布,超越初始阶段。
Approximate Gaussianity Beyond Initialisation in Neural Networks
- 用13参数置换不变高斯模型刻画权重矩阵相关性
- 发现该模型在训练全过程均有效,远超初始阶段
- 通过沃尔什距离量化分布演化,适合研究训练动态
针对MNIST分类任务,研究神经网络权重矩阵集合在训练过程中的分布特性,检验在高斯性和置换对称性假设下矩阵模型的有效性。发现13参数置换不变高斯矩阵模型能有效描述权重矩阵中的关联高斯性,其适用范围远超独立同分布高斯模型,且显著超出初始化阶段。表示论模型参数与图论特征共同提供了最优拟合模型的可解释框架,以及对偏离高斯性的微小变化的刻画。同时计算了该类模型的沃什距离,用于量化分布随训练的迁移。在整个研究中,测试了不同初始化方式、正则化、层深和层宽的影响,识别出特定偏差被增强的边界,并提出更通用但依然高度可解释的建模方法。
原文摘要 · Abstract (English)
Ensembles of neural network weight matrices are studied through the training process for the MNIST classification problem, testing the efficacy of matrix models for representing their distributions, under assumptions of Gaussianity and permutation-symmetry. The general 13-parameter permutation invariant Gaussian matrix models are found to be effective models for the correlated Gaussianity in the weight matrices, beyond the range of applicability of the simple Gaussian with independent identically distributed matrix variables, and notably well beyond the initialisation step. The representation theoretic model parameters, and the graph-theoretic characterisation of the permutation invariant matrix observables give an interpretable framework for the best-fit model and for small departures from Gaussianity. Additionally, the Wasserstein distance is calculated for this class of models and used to quantify the movement of the distributions over training. Throughout the work, the effects of varied initialisation regimes, regularisation, layer depth, and layer width are tested for this formalism, identifying limits where particular departures from Gaussianity are enhanced and how more general, yet still highly-interpretable, models can be developed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。