用小批量控制扩散动力学,高效采样神经网络参数空间。
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
- 用可控的小批量梯度噪声构造伪朗之万动力学,保持采样精度。
- 百万级参数网络下仍具高计算效率,优于传统混合蒙特卡洛方法。
- 中等温度采样可得类SGD泛化性能,无需验证集或早停。
根据玻尔兹曼分布对人工神经网络的参数空间进行采样,有助于理解低损失解的几何结构,并为训练提供一种替代传统损失最小化的方法。然而,如混合蒙特卡洛(hMC)这类精确采样方法在真实数据集上因需反复计算全批量梯度而变得计算成本过高。本文提出一种伪朗之万(pL)动力学,通过有控制地使用小批量,在大规模数据集上实现前馈神经网络的高效玻尔兹曼采样。该方法利用小批量梯度噪声的统计特性,调整虚构质量与摩擦系数,确保随机过程能有效采样目标平衡分布。数值验证显示,pL的动力学平衡统计与精确的hMC结果高度一致。性能基准测试表明,随着网络规模增大,hMC迅速失效,而pL方案保持高计算扩散性,可扩展至超过一百万参数的网络。最后,我们发现中等温度下的采样能获得与SGD相当的泛化性能,且无需验证集或早停机制。这些结果确立了受控小批量朗之万动力学作为探索和利用大型神经网络解空间的一种实用且可扩展的工具。
原文摘要 · Abstract (English)
Sampling the parameter space of artificial neural networks according to a Boltzmann distribution provides insight into the geometry of low-loss solutions and offers an alternative to conventional loss minimization for training. However, exact sampling methods such as hybrid Monte Carlo (hMC), while formally correct, become computationally prohibitive for realistic datasets because they require repeated evaluation of full-batch gradients. We introduce a pseudo-Langevin (pL) dynamics that enables efficient Boltzmann sampling of feed-forward neural networks trained with large datasets by using minibatches in a controlled manner. The method exploits the statistical properties of minibatch gradient noise and adjusts fictitious masses and friction coefficients to ensure that the induced stochastic process samples efficiently the desired equilibrium distribution. We validate numerically the approach by comparing its equilibrium statistics with those obtained from exact hMC sampling. Performance benchmarks demonstrate that, while hMC rapidly becomes inefficient as network size increases, the pL scheme maintains high computational diffusion and scales favorably to networks with over one million parameters. Finally, we show that sampling at intermediate temperatures yields optimal generalization performance, comparable to SGD, without requiring a validation set or early stopping procedure. These results establish controlled minibatch Langevin dynamics as a practical and scalable tool for exploring and exploiting the solution space of large neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。