提出非局部转移核,让RBM学习更高效稳定
Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines

- 设计跨RBM序列的往返式转移核,实现单步非局部采样
- 采样质量更高,所需迭代次数比BGS和DT减少30%以上
- 适合需要稳定训练的复杂RBM模型,尤其高能垒场景
学习受限玻尔兹曼机(RBMs)计算成本高,因其期望值通常难以精确计算。传统方法采用基于分块吉布斯采样(BGS)的局部马尔可夫链蒙特卡洛近似,但在存在高能量壁垒时采样质量差,影响学习效果。深度退火(DT)通过在一系列可学习的RBMs上并行退火缓解此问题,但需多步移动才能实现非局部转移。本文提出一种定义在DT所用RBMs序列上的转移核,具有序列往返结构,可在单次转移中实现非局部跳跃,同时保持序列不变。数值实验表明,该核能更频繁地执行非局部转移,以更少的转移次数获得更高采样质量;基于该核的学习更具稳定性,有效缓解了传统BGS与DT方法中的训练失败现象。
原文摘要 · Abstract (English)
Learning restricted Boltzmann machines (RBMs) is computationally challenging because it requires expectations whose exact evaluation is generally intractable. The expectations are typically evaluated using a sampling approximation based on blocked Gibbs sampling (BGS), which is a local Markov chain Monte Carlo transition kernel. However, the locality of BGS can lead to poor sampling quality when the RBM has high energy barriers, thereby degrading learning performance. Deep tempering (DT), which performs parallel tempering over a sequence of learnable RBMs including the training RBM, alleviates this locality issue. However, DT algorithmically requires multiple steps to move through the RBM sequence to achieve a nonlocal transition. In this paper, we propose a transition kernel defined over the RBM sequence used in DT. The proposed kernel has a round-trip structure over the sequence, enabling nonlocal moves within a single transition while leaving the RBM sequence invariant. Numerical experiments show that the proposed kernel performs nonlocal transitions more frequently and achieves higher sampling quality with fewer transitions than BGS and DT. We further verify that learning based on the proposed kernel is more stable and mitigates the training failures observed with BGS- and DT-based learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。