统一框架分析包含式KL散度的推断算法,提升变分推断效率。
Inclusive KL Gradient Flows: Otto-Wasserstein, Fisher-Rao-Gaussian, and Local-Estimator Dynamics
- 构建包含式KL散度的统一梯度流框架,涵盖多种几何结构。
- 在高斯分布流形上导出显式微分方程,支持高效变分推断。
- 提出局部估计器方法,无需密度比或核梯度,性能优于MMD方法。
Otto的Wasserstein梯度流为包含式(前向)Kullback-Leibler (KL) 散度提供了严谨的分析框架,但针对排除式(逆向)KL散度的算法较少使用此类工具。本文建立了一个统一的梯度流与概率密度函数框架,用于包含式KL推断。我们发现最大均值差异(MMD)最小化可视为使用近似梯度估计器的包含式KL推断,并发展了直接针对包含式KL散度的Fisher-Rao与Wasserstein-Fisher-Rao梯度流。将这些流限制在高斯分布流形上,可得到显式的梯度流常微分方程(ODE),为高斯变分推断奠定基础。基于此视角,进一步提出一种局部估计器的Wasserstein梯度流,其速度通过局部非参数回归获得,无需密度比评估或核梯度,相较于基于MMD的粒子方法提升了算法性能。
原文摘要 · Abstract (English)
Otto's Wasserstein gradient flow of the inclusive (forward) Kullback--Leibler (KL) divergence offers a principled framework for analyzing statistical inference algorithms, yet algorithms targeting the exclusive (reverse) KL divergence are rarely studied with such tools. We establish a unified gradient-flow and PDF framework for inclusive KL inference. We show that maximum mean discrepancy minimization can be viewed as inclusive KL inference with an approximate gradient estimator, and we develop the Fisher--Rao and Wasserstein--Fisher--Rao gradient flows that directly target the inclusive KL divergence. Restricting these flows to the manifold of Gaussian distributions yields explicit gradient-flow ODEs, providing a foundation for Gaussian variational inference. Building on this viewpoint, we further introduce a local-estimator Wasserstein gradient flow whose velocity is obtained by local nonparametric regression, free of density-ratio evaluation or kernel gradients, improving the algorithmic performance over the MMD-based particle method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。