arXiv:2410.22065stat.MLcs.LG2024-10NeurIPS被引 3

ReLU神经网络上哈密顿蒙特卡洛采样效率低下,因非可微性导致误差增大。

Hamiltonian Monte Carlo on ReLU Neural Networks is Inefficient

  • 利用跳跃积分器的哈密顿蒙特卡洛在ReLU网络中局部误差为Ω(ε)。
  • 相比经典O(ε³)误差,提出拒绝率更高,采样效率显著下降。
  • 理论分析与真实数据实验均验证了该方法在ReLU网络中的低效性。

我们分析了基于跳跃积分器的哈密顿蒙特卡洛算法在贝叶斯神经网络推断中的误差率。由于ReLU族激活函数的不可微性,使用这些激活函数的网络所对应的跳跃积分器哈密顿蒙特卡洛具有Ω(ε)的较大局部误差,而非经典的O(ε³)。这导致提议样本的拒绝率升高,使得该方法效率低下。我们通过经验模拟及真实数据集上的实验验证了上述理论结果,进一步凸显了在基于ReLU的神经网络上使用哈密顿蒙特卡洛进行推断的低效性。

原文摘要 · Abstract (English)

We analyze the error rates of the Hamiltonian Monte Carlo algorithm with leapfrog integrator for Bayesian neural network inference. We show that due to the non-differentiability of activation functions in the ReLU family, leapfrog HMC for networks with these activation functions has a large local error rate of $Ω(ε)$ rather than the classical error rate of $O(ε^3)$. This leads to a higher rejection rate of the proposals, making the method inefficient. We then verify our theoretical findings through empirical simulations as well as experiments on a real-world dataset that highlight the inefficiency of HMC inference on ReLU-based neural networks compared to analytical networks.

贝叶斯神经网络哈密顿蒙特卡洛非可微性推理效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。