arXiv:2410.05753stat.MLcs.LG2024-10被引 2

提出零方差控制变量法,显著降低变分推断中路径梯度的方差。

Pathwise Gradient Variance Reduction with Control Variates in Variational Inference

  • 引入零方差控制变量,无需复杂分布假设即可减少路径梯度方差。
  • 相比传统方法,新方法在简单和复杂变分族上均表现更优。
  • 适合需要高精度梯度估计的贝叶斯深度学习任务使用。

贝叶斯深度学习中的变分推断常需计算无闭式解的期望梯度。路径梯度估计器因方差较低而优于得分函数估计器,后者通常依赖方差缩减技术。然而,近期研究发现路径梯度也可受益于方差减少。本文回顾现有基于控制变量的路径梯度方差缩减方法,发现其多依赖积分近似,仅适用于简单变分族。为此,我们提出将零方差控制变量应用于路径梯度估计器,该方法仅需能从变分分布采样,对分布形式假设极少,具有更强通用性。

原文摘要 · Abstract (English)

Variational inference in Bayesian deep learning often involves computing the gradient of an expectation that lacks a closed-form solution. In these cases, pathwise and score-function gradient estimators are the most common approaches. The pathwise estimator is often favoured for its substantially lower variance compared to the score-function estimator, which typically requires variance reduction techniques. However, recent research suggests that even pathwise gradient estimators could benefit from variance reduction. In this work, we review existing control-variates-based variance reduction methods for pathwise gradient estimators to assess their effectiveness. Notably, these methods often rely on integrand approximations and are applicable only to simple variational families. To address this limitation, we propose applying zero-variance control variates to pathwise gradient estimators. This approach offers the advantage of requiring minimal assumptions about the variational distribution, other than being able to sample from it.

变分推断梯度估计方差缩减

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。