提出一种无需依赖维度的变分推断方法,显著提升高维问题求解效率。
Nearly Dimension-Independent Convergence of Mean-Field Black-Box Variational Inference
- 基于重参数化梯度的均值场变分推断,收敛速度近似与维度无关。
- 在强对数凹目标下,迭代次数仅随log d增长,远优于传统方法的O(d)。
- 适用于高维复杂分布推断,尤其适合对计算效率敏感的应用场景。
针对均值场位置尺度变分族,我们证明了使用重参数化梯度的黑箱变分推断(BBVI)在收敛速率上几乎不受维度影响。对于d维强对数凹且对数光滑的目标分布,采用子高斯族的BBVI达到ε-最优解所需的迭代次数仅具有O(log d)的维度依赖性,显著优于全秩位置尺度族的O(d)。对于重尾族,其维度依赖为O(d^{2/k}),其中k为族中有限矩的阶数。若目标对数密度的海森矩阵为常数,则复杂度完全无显式维度依赖。此外,我们还证明,决定该结果的关键梯度方差界无法仅通过目标对数密度海森矩阵的谱界进一步改进。
原文摘要 · Abstract (English)
We prove that, given a mean-field location-scale variational family, black-box variational inference (BBVI) with the reparametrization gradient converges at a rate that is nearly independent of explicit dimension dependence. Specifically, for a $d$-dimensional strongly log-concave and log-smooth target, the number of iterations for BBVI with a sub-Gaussian family to obtain a solution $ε$-close to the global optimum has a dimension dependence of $\mathrm{O}(\log d)$. This is a significant improvement over the $\mathrm{O}(d)$ dependence of full-rank location-scale families. For heavy-tailed families, we prove a weaker $\mathrm{O}(d^{2/k})$ dependence, where $k$ is the number of finite moments of the family. Additionally, if the Hessian of the target log-density is constant, the complexity is free of any explicit dimension dependence. We also prove that our bound on the gradient variance, which is key to our result, cannot be improved using only spectral bounds on the Hessian of the target log-density.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。