用粒子动力学训练隐变量能量模型,无需判别器且收敛性可保证。
Particle Dynamics for Latent-Variable Energy-Based Models
- 将最大似然训练转化为分布上的鞍点问题,通过耦合Wasserstein梯度流优化
- 在物理系统数值近似上表现优于同类方法,KL与W2距离均有收敛速率
- 理论严格更优的变分下界,适合追求可解释性与理论保障的研究者
隐变量能量模型(LVEBMs)为观测数据与隐变量的联合对分配单一归一化能量,兼具表达力与对隐藏结构的捕捉能力。本文将最大似然训练重新建模为在隐变量与联合流形上分布的鞍点问题,将内层更新视为耦合的Wasserstein梯度流。由此产生的算法在联合负池与条件隐变量粒子间交替进行阻尼Langevin更新及随机参数上升,无需判别器或辅助网络。在标准光滑性与耗散性假设下,证明了存在性与收敛性,给出KL散度与Wasserstein-2距离的衰减速率。鞍点视角还导出一个严格优于受限变分后验所获界限的ELBO。方法在物理系统数值近似任务上进行了评估,性能与现有方法相当。
原文摘要 · Abstract (English)
Latent-variable energy-based models (LVEBMs) assign a single normalized energy to joint pairs of observed data and latent variables, offering expressive generative modeling while capturing hidden structure. We recast maximum-likelihood training as a saddle problem over distributions on the latent and joint manifolds and view the inner updates as coupled Wasserstein gradient flows. The resulting algorithm alternates overdamped Langevin updates for a joint negative pool and for conditional latent particles with stochastic parameter ascent, requiring no discriminator or auxiliary networks. We prove existence and convergence under standard smoothness and dissipativity assumptions, with decay rates in KL divergence and Wasserstein-2 distance. The saddle-point view further yields an ELBO strictly tighter than bounds obtained with restricted amortized posteriors. Our method is evaluated on numerical approximations of physical systems and performs competitively against comparable approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。