用强化学习提升粒子采样效率,更好逼近复杂分布。
Reinforced sequential Monte Carlo for amortised sampling
- 将SMC与最大熵强化学习结合,用神经网络学采样策略和修正函数。
- 在多模态合成数据和丙氨酸二肽分布上,采样精度和训练稳定性均提升。
- 适合需要高效采样的高维复杂分布建模任务,如分子构象分析。
本文提出一种将可泛化采样与基于粒子的方法相结合的新框架,用于从未归一化的密度函数定义的分布中进行采样。我们建立了序列蒙特卡洛(SMC)与通过最大熵强化学习(MaxEnt RL)训练的神经序列采样器之间的联系,其中学习到的采样策略和价值函数分别定义了提议核和扭曲函数。利用这一联系,我们引入了一种离策略强化学习训练过程,使用SMC生成的样本作为行为策略,以更有效地探索目标分布。我们提出了稳定联合训练提议函数与扭曲函数的技术,并设计了一种自适应权重退火方案以降低训练信号方差。此外,借鉴经验回放思想,我们提出一种将历史样本与退火重要性采样权重结合的回放缓冲区机制。在连续与离散空间中的合成多模态目标以及丙氨酸二肽构象的玻尔兹曼分布上,实验表明该方法在逼近真实分布方面优于传统可泛化与蒙特卡洛方法,且训练更加稳定。
原文摘要 · Abstract (English)
This paper proposes a synergy of amortised and particle-based methods for sampling from distributions defined by unnormalised density functions. We state a connection between sequential Monte Carlo (SMC) and neural sequential samplers trained by maximum-entropy reinforcement learning (MaxEnt RL), wherein learnt sampling policies and value functions define proposal kernels and twist functions. Exploiting this connection, we introduce an off-policy RL training procedure for the sampler that uses samples from SMC -- using the learnt sampler as a proposal -- as a behaviour policy that better explores the target distribution. We describe techniques for stable joint training of proposals and twist functions and an adaptive weight tempering scheme to reduce training signal variance. Furthermore, building upon past attempts to use experience replay to guide the training of neural samplers, we derive a way to combine historical samples with annealed importance sampling weights within a replay buffer. On synthetic multi-modal targets (in both continuous and discrete spaces) and the Boltzmann distribution of alanine dipeptide conformations, we demonstrate improvements in approximating the true distribution as well as training stability compared to both amortised and Monte Carlo methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。