用可调控高阶统计量的数据模型,揭示神经网络学习顺序
Learning Beyond the Gaussian Data: Learning Dynamics of Neural Networks on an Expressive and Cumulant-Controllable Data Model
- 构建基于埃尔米特多项式的非高斯数据生成模型,可控偏度与峰度
- 实验显示网络先学均值协方差,再逐步学习高阶累积量
- 在Fashion-MNIST上验证,适合研究分布效应的理论分析
我们通过一个矩可控的非高斯数据模型,研究数据高阶统计特性对神经网络学习动态的影响。考虑到两层神经网络的表达能力,首先构建了一个生成式两层神经网络作为数据模型,其中激活函数通过埃尔米特多项式展开。这使得我们能够通过埃尔米特系数可解释地控制偏度、峰度等高阶累积量,同时保持数据模型的真实性。利用该模型生成的样本,我们对两层神经网络进行了受控在线学习实验。结果表明训练存在逐矩进展:网络先捕捉均值和协方差等低阶统计量,随后逐步学习高阶累积量。最后,我们在Fashion-MNIST数据集上预训练生成模型,并使用生成样本进行额外实验。结果确认了前述结论,并展示了该数据模型在真实场景中的实用性。总体而言,该方法弥合了简化数据假设与实际数据复杂性之间的差距,为机器学习与信号处理中分布效应的研究提供了原则性框架。
原文摘要 · Abstract (English)
We study the effect of high-order statistics of data on the learning dynamics of neural networks (NNs) by using a moment-controllable non-Gaussian data model. Considering the expressivity of two-layer neural networks, we first construct the data model as a generative two-layer NN where the activation function is expanded by using Hermite polynomials. This allows us to achieve interpretable control over high-order cumulants such as skewness and kurtosis through the Hermite coefficients while keeping the data model realistic. Using samples generated from the data model, we perform controlled online learning experiments with a two-layer NN. Our results reveal a moment-wise progression in training: networks first capture low-order statistics such as mean and covariance, and progressively learn high-order cumulants. Finally, we pretrain the generative model on the Fashion-MNIST dataset and leverage the generated samples for further experiments. The results of these additional experiments confirm our conclusions and show the utility of the data model in a real-world scenario. Overall, our proposed approach bridges simplified data assumptions and practical data complexity, which offers a principled framework for investigating distributional effects in machine learning and signal processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。