arXiv:2506.05088cs.LGstat.CO2025-06

提出新型核化估计器,让半隐式变分推断更稳定高效。

Semi-Implicit Variational Inference via Kernelized Path Gradient Descent

  • 用核平滑技术降低KL散度估计的方差和偏差。
  • 引入重要性采样修正,进一步减少函数空间中的偏差。
  • 理论证明与自洽斯坦因梯度下降等价,但梯度方差更低。

半隐式变分推断(SIVI)是逼近复杂后验分布的强大框架,但在高维场景下使用KL散度训练常面临方差大、偏差高的问题。尽管现有最优方法如核半隐式变分推断(KSIVI)在高维中表现良好,其训练仍较昂贵。本文提出一种核化KL散度估计器,通过非参数平滑实现训练稳定;为降低偏差,引入重要性采样修正。我们建立了与自洽斯坦因变分梯度下降的理论联系,发现两者最小化同一目标,但本方法梯度方差更低。此外,函数空间中的偏差具有良性质,带来更稳定的优化。实验表明,该方法在性能和训练效率上均优于或匹配当前最优SIVI方法。

原文摘要 · Abstract (English)

Semi-implicit variational inference (SIVI) is a powerful framework for approximating complex posterior distributions, but training with the Kullback-Leibler (KL) divergence can be challenging due to high variance and bias in high-dimensional settings. While current state-of-the-art semi-implicit variational inference methods, particularly Kernel Semi-Implicit Variational Inference (KSIVI), have been shown to work in high dimensions, training remains moderately expensive. In this work, we propose a kernelized KL divergence estimator that stabilizes training through nonparametric smoothing. To further reduce the bias, we introduce an importance sampling correction. We provide a theoretical connection to the amortized version of the Stein variational gradient descent, which estimates the score gradient via Stein's identity, showing that both methods minimize the same objective, but our semi-implicit approach achieves lower gradient variance. In addition, our method's bias in function space is benign, leading to more stable and efficient optimization. Empirical results demonstrate that our method outperforms or matches state-of-the-art SIVI methods in both performance and training efficiency.

变分推断核方法梯度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。