用非平衡统计力学方法改进能量模型训练,减少偏差,提升采样质量。
Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels
- 引入Jarzynski重加权理论,优化能量模型的训练过程。
- 在扩散模型中降低离散化误差,提升生成样本质量。
- 适用于需高精度采样的生成建模任务,如复杂分布建模。
能量模型(EBM)为生成建模提供了灵活框架,但其训练在理论上仍具挑战性,主要源于归一化常数的近似与从复杂多模态分布中高效采样的困难。传统方法如对比散度和得分匹配会引入偏差,影响学习准确性。本文对来自非平衡统计力学的Jarzynski重加权技术进行理论分析,重点探讨过渡核选择的影响,并将其应用于两类生成框架:(i) 基于流的扩散模型,通过随机插值重新解释Jarzynski重加权,以缓解离散化误差并改善样本质量;(ii) 受限玻尔兹曼机,分析其在修正对比散度偏差中的作用。结果揭示了核选择与模型性能之间的内在关联,凸显了Jarzynski重加权作为生成学习中可信赖工具的潜力。
原文摘要 · Abstract (English)
Energy-Based Models (EBMs) provide a flexible framework for generative modeling, but their training remains theoretically challenging due to the need to approximate normalization constants and efficiently sample from complex, multi-modal distributions. Traditional methods, such as contrastive divergence and score matching, introduce biases that can hinder accurate learning. In this work, we present a theoretical analysis of Jarzynski reweighting, a technique from non-equilibrium statistical mechanics, and its implications for training EBMs. We focus on the role of the choice of the kernel and we illustrate these theoretical considerations in two key generative frameworks: (i) flow-based diffusion models, where we reinterpret Jarzynski reweighting in the context of stochastic interpolants to mitigate discretization errors and improve sample quality, and (ii) Restricted Boltzmann Machines, where we analyze its role in correcting the biases of contrastive divergence. Our results provide insights into the interplay between kernel choice and model performance, highlighting the potential of Jarzynski reweighting as a principled tool for generative learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。