Langevin算法可无条件学习任意规模的两层神经网络,且收敛速度可证明。
Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
- 用光滑激活函数和弗罗贝尼乌斯正则化,证明算法在任意网络规模下收敛。
- 在任意数据上,迭代序列以非渐近速率逼近吉布斯分布,正则化量与网络大小无关。
- 适合关注理论深度、优化与泛化关系的研究者,尤其对统计学习基础感兴趣者。
本文证明,在任意规模的两层神经网络和任意数据集上,拉普拉斯蒙特卡洛(Langevin Monte-Carlo)算法均可实现学习,并给出非渐近收敛速率。通过证明在 q-Renyi 散度下,该算法的迭代点列会收敛至弗罗贝尼乌斯范数正则化损失对应的吉布斯分布,我们建立了这一结果。关键在于,为保证结论成立所需的正则化量与网络规模无关。该成果整合了近期关于等周条件与LMC收敛性、以及两层神经网络损失函数可通过恒定正则化满足Villani条件的观察,从而确保其吉布斯测度满足Poincaré不等式。
原文摘要 · Abstract (English)
In this work, we will establish that the Langevin Monte-Carlo algorithm can learn depth-2 neural nets of any size and for any data and we give non-asymptotic convergence rates for it. We achieve this via showing that in q-Renyi divergence, the iterates of Langevin Monte Carlo converge to the Gibbs distribution of Frobenius norm regularized losses for any of these nets, when using smooth activations and in both classification and regression settings. Most critically, the amount of regularization needed for our results is independent of the size of the net. This result achieves a synthesis of several recent observations about isoperimetry conditions under which LMC converges and that two-layer neural loss functions can always be regularized by a certain constant amount such that they satisfy the Villani conditions, and thus their Gibbs measures satisfy a Poincare inequality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。