用信息论方法量化神经网络输出趋近高斯分布的速度。
Entropic bounds for conditionally Gaussian vectors and applications to neural networks
- 基于信息熵不等式,给出条件高斯与普通高斯的距离上界。
- 证明随机初始化的全连接网络在层宽无限时以最优速率收敛至高斯分布。
- 适用于多种激活函数,可推广到贝叶斯后验分布的精确估计。
利用信息论中的熵不等式,我们为条件高斯分布与可逆协方差矩阵的高斯分布之间的总变差距离和2-Wasserstein距离提供了新的上界。将这些结果应用于随机初始化的全连接神经网络及其导数在有限输入点上的行为:当初始分布为高斯且内部层大小趋于无穷时,我们量化了其收敛到高斯分布的速度。该分析对激活函数假设较弱,可恢复多种距离下的最优收敛速率,从而改进并拓展了Basteri and Trevisan (2023)、Favaro et al. (2023)、Trevisan (2024) 和 Apollonio et al. (2024) 的成果。主要工具是Hanin (2024) 建立的定量累积量估计。作为应用,我们给出了神经网络及其导数的贝叶斯后验分布与对应高斯极限之间的总变差距离上界,得到了Hron et al. (2022) 后验中心极限定理的定量版本,并将Trevisan (2024) 的若干估计扩展至总变差度量。
原文摘要 · Abstract (English)
Using entropic inequalities from information theory, we provide new bounds on the total variation and 2-Wasserstein distances between a conditionally Gaussian law and a Gaussian law with invertible covariance matrix. We apply our results to quantify the speed of convergence to Gaussian of a randomly initialized fully connected neural network and its derivatives - evaluated in a finite number of inputs - when the initialization is Gaussian and the sizes of the inner layers diverge to infinity. Our results require mild assumptions on the activation function, and allow one to recover optimal rates of convergence in a variety of distances, thus improving and extending the findings of Basteri and Trevisan (2023), Favaro et al. (2023), Trevisan (2024) and Apollonio et al. (2024). One of our main tools are the quantitative cumulant estimates established in Hanin (2024). As an illustration, we apply our results to bound the total variation distance between the Bayesian posterior law of the neural network and its derivatives, and the posterior law of the corresponding Gaussian limit: this yields quantitative versions of a posterior CLT by Hron et al. (2022), and extends several estimates by Trevisan (2024) to the total variation metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。