改进隐层激活方式可显著提升受限玻尔兹曼机的存储与召回能力。
On the role of non-linear latent features in bipartite generative neural networks
- 通过引入更丰富的隐层先验分布,优化模型能量特性。
- 在有限温度下,新设计使回忆成功率大幅提升,突破传统二值隐层限制。
- 理论分析结合模拟验证,适合研究神经网络记忆机制的研究者。
我们研究了双分图能量基神经网络(即受限玻尔兹曼机,RBMs)的相图与记忆检索能力,考察其隐藏单元的先验分布(包括二值、多状态及类似ReLU的激活函数)的影响。借鉴霍普菲尔德模型,并运用无序系统统计物理的解析工具,分析架构选择与激活函数如何塑造模型热力学性质。结果表明,采用二值隐藏节点且连接密度高的标准RBMs存在临界容量降低问题,制约其作为关联记忆的有效性。为解决此问题,我们探讨引入局部偏置和更丰富的隐藏单元先验等修改方案。这些调整恢复了有序检索相,显著提升有限温度下的召回性能。理论结果经有限尺寸蒙特卡洛模拟验证,凸显隐藏单元设计对增强RBMs表达能力的关键作用。
原文摘要 · Abstract (English)
We investigate the phase diagram and memory retrieval capabilities of bipartite energy-based neural networks, namely Restricted Boltzmann Machines (RBMs), as a function of the prior distribution imposed on their hidden units - including binary, multi-state, and ReLU-like activations. Drawing connections to the Hopfield model and employing analytical tools from statistical physics of disordered systems, we explore how the architectural choices and activation functions shape the thermodynamic properties of these models. Our analysis reveals that standard RBMs with binary hidden nodes and extensive connectivity suffer from reduced critical capacity, limiting their effectiveness as associative memories. To address this, we examine several modifications, such as introducing local biases and adopting richer hidden unit priors. These adjustments restore ordered retrieval phases and markedly improve recall performance, even at finite temperatures. Our theoretical findings, supported by finite-size Monte Carlo simulations, highlight the importance of hidden unit design in enhancing the expressive power of RBMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。