提出低方差估计器,显著提升贝叶斯推断在布雷斯-瓦瑟斯坦流形上的优化效率。
Stochastic variance-reduced Gaussian variational inference on the Bures-Wasserstein manifold
- 基于控制变量原理设计新型方差缩减估计器
- 相比传统方法性能提升数量级,理论证明方差更小
- 适合需要高精度的贝叶斯推断与变分推断场景
在机器学习领域,布雷斯-瓦瑟斯坦空间中的优化日益受到关注,因其将变分推断与瓦瑟斯坦梯度流联系起来。变分推断的目标函数(基于KL散度)可表示为负熵与势能之和,使前向-后向欧拉法成为首选方案。值得注意的是,后向步具有闭式解,提升了实用性;但前向步不精确,因势能的布雷斯-瓦瑟斯坦梯度涉及‘不可计算’期望。现有方法通常使用蒙特卡洛近似(实践中常为单样本估计),导致方差高、性能差。本文提出一种基于控制变量原理的新方差缩减估计器。理论上证明该估计器在相关场景下方差小于蒙特卡洛估计器,并证明方差缩减有助于改进当前分析中的优化界。实验表明,该方法相较以往布雷斯-瓦瑟斯坦方法实现数量级性能提升。
原文摘要 · Abstract (English)
Optimization in the Bures-Wasserstein space has been gaining popularity in the machine learning community since it draws connections between variational inference and Wasserstein gradient flows. The variational inference objective function of Kullback-Leibler divergence can be written as the sum of the negative entropy and the potential energy, making forward-backward Euler the method of choice. Notably, the backward step admits a closed-form solution in this case, facilitating the practicality of the scheme. However, the forward step is not exact since the Bures-Wasserstein gradient of the potential energy involves "intractable" expectations. Recent approaches propose using the Monte Carlo method -- in practice a single-sample estimator -- to approximate these terms, resulting in high variance and poor performance. We propose a novel variance-reduced estimator based on the principle of control variates. We theoretically show that this estimator has a smaller variance than the Monte-Carlo estimator in scenarios of interest. We also prove that variance reduction helps improve the optimization bounds of the current analysis. We demonstrate that the proposed estimator gains order-of-magnitude improvements over the previous Bures-Wasserstein methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。