提出无需模拟的神经采样训练方法,揭示预处理对避免模式崩溃的关键作用。
No Trick, No Treat: Pursuits and Challenges Towards Simulation-free Training of Neural Samplers
- 用时变归一化流实现无模拟训练,简化了传统采样流程
- 缺乏Langevin预处理时,多数方法在简单目标上仍严重模式崩溃
- 结合并行退火与生成模型,提供强基线以指导未来研究
我们研究采样问题,即从密度仅知归一化常数的分布中抽取样本。近年来生成建模在逼近高维数据分布方面取得突破,推动了基于神经网络的采样方法发展。然而,神经采样器通常因训练中需模拟轨迹而带来巨大计算开销。这促使人们探索无模拟训练方法。本文提出对已有方法的优雅改进,借助时变归一化流实现无模拟训练,但最终仍出现严重模式崩溃。深入分析发现,几乎所有成功的神经采样器都依赖Langevin预处理来避免模式崩溃。我们系统比较了多种主流方法及其目标函数,证明在缺乏Langevin预处理的情况下,多数方法无法充分覆盖简单目标分布。最后,我们提出一种强基线:将最先进的MCMC方法并行退火(Parallel Tempering, PT)与生成模型结合,为未来神经采样器的研究提供新方向。
原文摘要 · Abstract (English)
We consider the sampling problem, where the aim is to draw samples from a distribution whose density is known only up to a normalization constant. Recent breakthroughs in generative modeling to approximate a high-dimensional data distribution have sparked significant interest in developing neural network-based methods for this challenging problem. However, neural samplers typically incur heavy computational overhead due to simulating trajectories during training. This motivates the pursuit of simulation-free training procedures of neural samplers. In this work, we propose an elegant modification to previous methods, which allows simulation-free training with the help of a time-dependent normalizing flow. However, it ultimately suffers from severe mode collapse. On closer inspection, we find that nearly all successful neural samplers rely on Langevin preconditioning to avoid mode collapsing. We systematically analyze several popular methods with various objective functions and demonstrate that, in the absence of Langevin preconditioning, most of them fail to adequately cover even a simple target. Finally, we draw attention to a strong baseline by combining the state-of-the-art MCMC method, Parallel Tempering (PT), with an additional generative model to shed light on future explorations of neural samplers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。