提出自适应方差采样法,加速多模态分布采样。
Sampling with Adaptive Variance for Multimodal Distributions
- 基于自适应扩散系数的动态系统,可看作加权Wasserstein梯度流。
- 在非凸势能下收敛速度显著快于经典朗之万动力学。
- 无需梯度信息,适合高维复杂分布采样,尤其适用于多峰问题。
我们提出并分析了一类针对有界域上多模态分布的自适应采样算法,其结构与经典的阻尼朗之万动力学相似。首先证明这类具有自适应扩散系数和向量场的线性动力学可被解释为当前分布与目标吉布斯分布之间相对熵(KL散度)的加权Wasserstein梯度流,直接导致KL和χ²散度的指数收敛,收敛速率依赖于加权Wasserstein度量和吉布斯势能。随后表明,该动力学的无导数版本可在不依赖吉布斯势能梯度信息的情况下用于采样;对于具有非凸势能的吉布斯分布,该方法可实现比经典阻尼朗之万动力学更快的收敛速度。对非凸势能中局部极小值间平均转移时间的比较进一步凸显了无导数动力学在采样上的更高效率。
原文摘要 · Abstract (English)
We propose and analyze a class of adaptive sampling algorithms for multimodal distributions on a bounded domain, which share a structural resemblance to the classic overdamped Langevin dynamics. We first demonstrate that this class of linear dynamics with adaptive diffusion coefficients and vector fields can be interpreted and analyzed as weighted Wasserstein gradient flows of the Kullback--Leibler (KL) divergence between the current distribution and the target Gibbs distribution, which directly leads to the exponential convergence of both the KL and $χ^2$ divergences, with rates depending on the weighted Wasserstein metric and the Gibbs potential. We then show that a derivative-free version of the dynamics can be used for sampling without gradient information of the Gibbs potential and that for Gibbs distributions with nonconvex potentials, this approach could achieve significantly faster convergence than the classical overdamped Langevin dynamics. A comparison of the mean transition times between local minima of a nonconvex potential further highlights the better efficiency of the derivative-free dynamics in sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。