用沃瑟斯坦几何优化变分推断,让采样越多越稳定高效。
Bures-Wasserstein Importance-Weighted Evidence Lower Bound: Exposition and Applications
- 在沃瑟斯坦流形上重定义变分下界,改进梯度估计。
- 采样数增大时梯度信噪比仍保持√K增长,优于传统方法。
- 适合需要高精度近似和大样本采样的贝叶斯推断任务。
重要性加权证据下界(IW-ELBO)是变分推断中一种有效目标,可收紧标准ELBO并缓解模式聚焦问题。然而,在欧几里得空间优化IW-ELBO常因梯度估计器信噪比(SNR)消失而效率低下。本文将IW-ELBO的优化置于玻尔斯-沃瑟斯坦空间(即高斯分布流形上的2-Wasserstein度量),推导了其沃瑟斯坦梯度,并投影至该空间以获得适用于高斯变分推断的可计算算法。分析的关键贡献在于梯度估计器的稳定性:尽管欧几里得梯度的信噪比随重要性样本数K增加而趋于消失,我们证明沃瑟斯坦梯度的信噪比可保证以Ω(√K)速度增长,确保在大K下仍具优化效率。进一步将此几何分析扩展至变分瑞尼重要性加权自编码器(VR-ELBO),建立了类似稳定性保证。实验表明,所提框架相比其他基线实现了更优的近似性能。
原文摘要 · Abstract (English)
The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour. However, optimizing the IW-ELBO in Euclidean space is often inefficient, as its gradient estimators suffer from a vanishing signal-to-noise ratio (SNR). This paper formulates the optimisation of the IW-ELBO in Bures-Wasserstein space, a manifold of Gaussian distributions equipped with the 2-Wasserstein metric. We derive the Wasserstein gradient of the IW-ELBO and project it onto the Bures-Wasserstein space to yield a tractable algorithm for Gaussian VI. A pivotal contribution of our analysis concerns the stability of the gradient estimator. While the SNR of the standard Euclidean gradient estimator is known to vanish as the number of importance samples $K$ increases, we prove that the SNR of the Wasserstein gradient scales favourably as $Ω(\sqrt{K})$, ensuring optimisation efficiency even for large $K$. We further extend this geometric analysis to the Variational Rényi Importance-Weighted Autoencoder bound, establishing analogous stability guarantees. Experiments demonstrate that the proposed framework achieves superior approximation performance compared to other baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。