用几何方法分析变分推断收敛性,突破传统瓶颈。
Geometric Convergence Analysis of Variational Inference via Bregman Divergences
- 将负ELBO重构为Bregman散度,利用指数族结构建立几何视角。
- 证明梯度下降在固定与递减步长下均有非渐近收敛率。
- 适合研究推断理论或优化算法的学者,尤其关注收敛性分析者。
变分推断(VI)通过优化证据下界(ELBO)提供可扩展的贝叶斯推断框架,但由于目标函数在欧氏空间中的非凸性和非光滑性,其收敛性分析仍具挑战。本文利用分布的指数族结构,提出一种新的理论框架:将负ELBO表示为关于对数归一化函数的Bregman散度,从而实现对优化景观的几何分析。我们证明该Bregman表示具有弱单调性,虽弱于凸性,但足以支持严格的收敛性分析。通过推导参数空间射线上的目标函数界,揭示了由费舍尔信息矩阵谱特性决定的性质。在此几何框架下,我们证明了梯度下降算法在常数和递减步长下的非渐近收敛速率。
原文摘要 · Abstract (English)
Variational Inference (VI) provides a scalable framework for Bayesian inference by optimizing the Evidence Lower Bound (ELBO), but convergence analysis remains challenging due to the objective's non-convexity and non-smoothness in Euclidean space. We establish a novel theoretical framework for analyzing VI convergence by exploiting the exponential family structure of distributions. We express negative ELBO as a Bregman divergence with respect to the log-partition function, enabling a geometric analysis of the optimization landscape. We show that this Bregman representation admits a weak monotonicity property that, while weaker than convexity, provides sufficient structure for rigorous convergence analysis. By deriving bounds on the objective function along rays in parameter space, we establish properties governed by the spectral characteristics of the Fisher information matrix. Under this geometric framework, we prove non-asymptotic convergence rates for gradient descent algorithms with both constant and diminishing step sizes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。