提出自调节温度机制,解决有限步吉布斯采样中的模型失稳问题。
Thermodynamic Regulation of Finite-Time Gibbs Training in Energy-Based Models: A Restricted Boltzmann Machine Study
- 引入随采样统计量动态变化的温度变量,实现内生热力学调控
- 实验显示自调节模型在MNIST上样本多样性提升,参数更稳定
- 适合研究能量模型训练不稳定性或需高采样质量的场景
受限玻尔兹曼机(RBMs)通常在固定采样温度下使用有限长度的吉布斯链进行训练。这一做法隐含假设:随着学习过程能量景观演化,随机性仍保持有效。然而,在非凸能量模型中,固定温度下的有限时间训练可能导致有效场增强与电导塌陷,使吉布斯采样渐进冻结、负相局部化,且在缺乏足够正则化时参数出现确定性线性漂移。为解决此不稳定性,本文提出一种内生热力学调控框架,让温度作为耦合采样统计量的动力学状态变量。在标准局部利普希茨条件与双时间尺度分离假设下,证明了严格正L2正则化下参数全局有界性。进一步证明热力学子系统局部指数稳定,并表明调控区间可缓解逆温度爆炸与冻结引发的退化现象。在MNIST上的实验表明,所提自调节RBM显著提升归一化稳定性与有效样本规模,同时保持重建性能。总体而言,结果将RBM训练重新诠释为受控的非平衡动力学过程,而非静态平衡近似。
原文摘要 · Abstract (English)
Restricted Boltzmann Machines (RBMs) are typically trained using finite-length Gibbs chains under a fixed sampling temperature. This practice implicitly assumes that the stochastic regime remains valid as the energy landscape evolves during learning. We argue that this assumption can become structurally fragile under finite-time training dynamics. This fragility arises because, in nonconvex energy-based models, fixed-temperature finite-time training can generate admissible trajectories with effective-field amplification and conductance collapse. As a result, the Gibbs sampler may asymptotically freeze, the negative phase may localize, and, without sufficiently strong regularization, parameters may exhibit deterministic linear drift. To address this instability, we introduce an endogenous thermodynamic regulation framework in which temperature evolves as a dynamical state variable coupled to measurable sampling statistics. Under standard local Lipschitz conditions and a two-time-scale separation regime, we establish global parameter boundedness under strictly positive L2 regularization. We further prove local exponential stability of the thermodynamic subsystem and show that the regulated regime mitigates inverse-temperature blow-up and freezing-induced degeneracy within a forward-invariant neighborhood. Experiments on MNIST demonstrate that the proposed self-regulated RBM substantially improves normalization stability and effective sample size relative to fixed-temperature baselines, while preserving reconstruction performance. Overall, the results reinterpret RBM training as a controlled non-equilibrium dynamical process rather than a static equilibrium approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。