提出SVGD的粒子收敛率新方法,实现近最优精度提升
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent
- 通过分解相对熵导数,建立基于核Stein散度的收敛分析框架
- 在连续与离散时间下均获1/√N阶收敛率,优于此前结果
- 适用于高维场景,适合从事概率推断与采样算法研究者
本文给出了斯坦因变分梯度下降(SVGD)算法在核斯坦因散度(KSD)和Wasserstein-2度量下的有限粒子收敛速率。核心思路是:从正则初始分布出发,N个粒子联合密度与N重目标测度间相对熵的时间导数可分解为一个与N倍期望KSD²成正比的主导负项,以及一个较小的正项。该分解导出在连续与离散时间下均为1/√N阶的KSD收敛率,实现对Shi和Mackey(2024)结果的近乎最优双指数改进。在核函数与势函数的温和假设下,这些界随维度d多项式增长。通过引入双线性核项,进一步在连续时间下获得Wasserstein-2收敛性。对于'双线性+Matérn'核,得到具有类似独立同分布设置下维度灾难特性的Wasserstein-2速率。还获得了时间平均粒子律的边际收敛与长时间混沌传播结果。
原文摘要 · Abstract (English)
We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\mathsf{KSD}$) and Wasserstein-2 metrics. Our key insight is that the time derivative of the relative entropy between the joint density of $N$ particle locations and the $N$-fold product target measure, starting from a regular initial distribution, splits into a dominant `negative part' proportional to $N$ times the expected $\mathsf{KSD}^2$ and a smaller `positive part'. This observation leads to $\mathsf{KSD}$ rates of order $1/\sqrt{N}$, in both continuous and discrete time, providing a near optimal (in the sense of matching the corresponding i.i.d. rates) double exponential improvement over the recent result by Shi and Mackey (2024). Under mild assumptions on the kernel and potential, these bounds also grow polynomially in the dimension $d$. By adding a bilinear component to the kernel, the above approach is used to further obtain Wasserstein-2 convergence in continuous time. For the case of `bilinear + Matérn' kernels, we derive Wasserstein-2 rates that exhibit a curse-of-dimensionality similar to the i.i.d. setting. We also obtain marginal convergence and long-time propagation of chaos results for the time-averaged particle laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。