证明了贝叶斯PCA的CAVI算法呈指数收敛,首次给出理论保证。
Exponential Convergence of CAVI for Bayesian PCA
- 通过连接幂迭代法,推导出单主成分时的精确指数收敛性。
- 推广至多主成分情形,证明了同样指数收敛,但形式略有不同。
- 提出新的多元正态分布KL散度下界,对信息论有独立价值。
概率主成分分析(PCA)及其贝叶斯变体(BPCA)在机器学习与统计学中广泛用于降维。相比传统方法,BPCA的优势在于可量化不确定性。参数通常通过均值场变分推断学习,特别是坐标上升变分推断(CAVI)算法。然而,迄今尚未明确CAVI在BPCA中的收敛速度。本文填补了这一空白:首先,在仅使用一个主成分的情况下,通过与经典幂迭代算法的联系,证明了精确的指数收敛性,表明传统PCA可作为BPCA参数的点估计;其次,借助最新工具,证明了任意数量主成分情形下CAVI的指数收敛性,得出更普遍的结果,但形式略有差异。为此,我们引入了一个关于两个多元正态分布对称KL散度的新下界,该结果或对信息论具有独立意义。
原文摘要 · Abstract (English)
Probabilistic principal component analysis (PCA) and its Bayesian variant (BPCA) are widely used for dimension reduction in machine learning and statistics. The main advantage of probabilistic PCA over the traditional formulation is allowing uncertainty quantification. The parameters of BPCA are typically learned using mean-field variational inference, and in particular, the coordinate ascent variational inference (CAVI) algorithm. So far, the convergence speed of CAVI for BPCA has not been characterized. In our paper, we fill this gap in the literature. Firstly, we prove a precise exponential convergence result in the case where the model uses a single principal component (PC). Interestingly, this result is established through a connection with the classical $\textit{power iteration algorithm}$ and it indicates that traditional PCA is retrieved as points estimates of the BPCA parameters. Secondly, we leverage recent tools to prove exponential convergence of CAVI for the model with any number of PCs, thus leading to a more general result, but one that is of a slightly different flavor. To prove the latter result, we additionally needed to introduce a novel lower bound for the symmetric Kullback--Leibler divergence between two multivariate normal distributions, which, we believe, is of independent interest in information theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。