研究高维下RBM的训练动态,揭示其达到最优弱恢复极限。
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
- 将RBM简化为非可分正则化的多指标模型,便于分析
- 在有突变协方差数据上,训练达到BBP相变阈值
- 适用于理解高维无监督学习中的优化行为
受限玻尔兹曼机(RBM)是最简单的生成神经网络之一,能够学习输入分布。尽管结构简单,其在从训练数据中学习的表现仅在本质上等同于数据奇异值分解的情况下被充分理解。本文考虑输入空间维度很大而隐层单元数恒定的极限情形。在此极限下,我们将标准RBM训练目标简化为等价于具有非可分正则化的多指标模型的形式,从而可应用针对多指标模型已建立的方法进行分析,如近似消息传递(AMP)及其状态演化、以及通过动力学平均场理论对梯度下降(GD)的分析。我们进一步给出了在由突变协方差模型生成的数据上的训练动态的严格渐近结果,该模型是适合无监督学习的结构的原型。特别地,我们证明了RBM在突变协方差模型中达到了最优计算弱恢复阈值,与BBP相变一致。
原文摘要 · Abstract (English)
The Restricted Boltzmann Machine (RBM) is one of the simplest generative neural networks capable of learning input distributions. Despite its simplicity, the analysis of its performance in learning from the training data is only well understood in cases that essentially reduce to singular value decomposition of the data. Here, we consider the limit of a large dimension of the input space and a constant number of hidden units. In this limit, we simplify the standard RBM training objective into a form that is equivalent to the multi-index model with non-separable regularization. This opens a path to analyze training of the RBM using methods that are established for multi-index models, such as Approximate Message Passing (AMP) and its state evolution, and the analysis of Gradient Descent (GD) via the dynamical mean-field theory. We then give rigorous asymptotics of the training dynamics of RBM on data generated by the spiked covariance model as a prototype of a structure suitable for unsupervised learning. We show in particular that RBM reaches the optimal computational weak recovery threshold, aligning with the BBP transition, in the spiked covariance model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。