提出一种在函数空间中近似贝叶斯后验的高效算法,可逼近最优变分解。
Near-Optimal Approximations for Bayesian Inference in Function Space
- 通过截断K-L展开将无限维扩散过程投影到有限维空间进行近似。
- 计算复杂度为O(M³ + JM²),在凸和Lipschitz连续负对数似然下逼近最优变分解。
- 适用于需要高精度非参数变分推断的研究者,尤其在函数空间建模中优势明显。
我们提出一种可扩展的贝叶斯后验推理算法,定义在再生核希尔伯特空间(RKHS)上。给定似然函数和表示先验的高斯随机元,对应的贝叶斯后验测度Π_{B}可作为RKHS值朗之万扩散的平稳分布获得。通过将无穷维朗之万扩散投影到Kosambi-Karhunen-Loève展开的前M个分量,得到这M个分量的近似后验,并基于全概率定律与充分性假设完成Π_{B}的推断。该方法计算复杂度为O(M³ + JM²),其中J为从后验测度Π_{B}生成的样本数。有趣的是,该算法在特定情况下恢复了稀疏变分高斯过程(SVGP,Titsias, 2009)的后验,因两者均依赖于充分性假设。然而,不同于将后验限制为高斯过程的参数化方法,我们的方法基于非参数变分族𝒫(ℝ^M),即ℝ^M上的所有概率测度。因此,在负对数似然为凸且Lipschitz连续时,其结果在𝒫(ℝ^M)中逼近最优的M维变分近似,且在高斯误差似然下与SVGP完全一致。
原文摘要 · Abstract (English)
We propose a scalable inference algorithm for Bayes posteriors defined on a reproducing kernel Hilbert space (RKHS). Given a likelihood function and a Gaussian random element representing the prior, the corresponding Bayes posterior measure $Π_{\text{B}}$ can be obtained as the stationary distribution of an RKHS-valued Langevin diffusion. We approximate the infinite-dimensional Langevin diffusion via a projection onto the first $M$ components of the Kosambi-Karhunen-Loève expansion. Exploiting the thus obtained approximate posterior for these $M$ components, we perform inference for $Π_{\text{B}}$ by relying on the law of total probability and a sufficiency assumption. The resulting method scales as $O(M^3+JM^2)$, where $J$ is the number of samples produced from the posterior measure $Π_{\text{B}}$. Interestingly, the algorithm recovers the posterior arising from the sparse variational Gaussian process (SVGP) (see Titsias, 2009) as a special case, owed to the fact that the sufficiency assumption underlies both methods. However, whereas the SVGP is parametrically constrained to be a Gaussian process, our method is based on a non-parametric variational family $\mathcal{P}(\mathbb{R}^M)$ consisting of all probability measures on $\mathbb{R}^M$. As a result, our method is provably close to the optimal $M$-dimensional variational approximation of the Bayes posterior $Π_{\text{B}}$ in $\mathcal{P}(\mathbb{R}^M)$ for convex and Lipschitz continuous negative log likelihoods, and coincides with SVGP for the special case of a Gaussian error likelihood.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。