arXiv:2410.12035stat.MLcs.LG2024-10

对比重要性加权变分推断中两种梯度估计器,发现双重重参数化更优。

Learning with Importance Weighted Variational Inference

  • 统一分析重参数化与双重重参数化梯度估计器的性能差异。
  • 理论证明双重重参数化在采样数增大时方差更小,收敛更稳定。
  • 适用于高维复杂模型的变分推断优化,尤其适合初始值较差的情况。

涉及重要性加权思想的若干变分界(如IWAE、VR、VR-IWAE)推广了边缘似然优化中的证据下界(ELBO)。然而,边界选择与梯度估计器联合使用时对变分推断算法行为的影响仍不明确。本文对基于IWAE、VR和VR-IWAE边界的重参数化(REP)与双重重参数化(DREP)梯度估计器进行了统一理论比较。通过分析蒙特卡洛样本数 $N \to \infty$ 时信噪比的渐近性质,揭示了梯度估计器中的偏差-方差权衡,并正式证明了在重要性加权变分推断中DREP优于REP。进一步对极端情形($N$ 和变分分布与后验分布之间的KL散度均趋于无穷)的渐近分析表明,即使变分近似质量下降,重要性加权的梯度估计仍指向合理方向。这些互补结果刻画了从初始不佳到最终收敛的优化轨迹。此外,本文的证明方法为样本均值比的理论研究提供了通用工具,其影响超出变分推断范畴,是蒙特卡洛方法领域的独立贡献。

原文摘要 · Abstract (English)

Several variational bounds involving importance weighting ideas generalize the Evidence Lower BOund (ELBO) for marginal likelihood optimization, such as the Importance-weighted Auto-Encoder (IWAE), Variational Rényi (VR) and VR-IWAE bounds. Yet, it remains unclear how the joint choice of bound and gradient estimator impacts the behavior of the resulting variational inference (VI) algorithms. This paper provides a unified theoretical comparison of reparameterized (REP) and doubly-reparameterized (DREP) gradient estimators tied to the IWAE, VR and VR-IWAE bounds. Through asymptotic analyses of the Signal-to-Noise Ratio as the number of Monter Carlo samples $N$ goes to infinity, we identify a bias-variance tradeoff in these gradient estimators and we formally justify the superiority of DREP over REP in importance-weighted VI. An additional asymptotic analysis for challenging regimes, where both $N$ and the Kullback-Leibler divergence between the variational and posterior densities go to infinity, indicates that importance-weighted VI gradient estimators point in a well-founded direction even when the variational approximation deteriorates. Together, these complementary results characterize the optimization trajectory in importance-weighted VI from poor initialization to final convergence. Importantly, our proof techniques establish general theoretical tools for the study of sample means ratios whose scope extend beyond VI and constitute an independent contribution to the field of Monte Carlo methods.

变分推断重要性加权梯度估计蒙特卡洛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。