arXiv:2510.21598stat.MLcs.LG2025-10NeurIPS被引 1

用多重t分布乘积建模复杂分布,实现高效变分推断。

Fisher meets Feynman: score-based variational inference with a product of experts

  • 将t分布加权乘积作为变分族,支持多峰、长尾和偏斜分布建模。
  • 通过狄利克雷潜变量重写乘积形式,实现从复杂分布中加权采样。
  • 基于分数匹配的迭代优化收敛快,适用于真实数据与合成数据。

我们提出一种高度表达且计算可处理的黑箱变分推断(BBVI)族。该族中的每个成员为加权专家乘积(PoE),每个加权专家正比于多元t分布。这些乘积能捕捉具有偏度、重尾和多峰特征的分布,但需能从中采样才能用于BBVI。我们通过引入辅助狄利克雷随机变量,将乘积重新表述为潜变量模型,该方法源自费曼在量子场论中用于环积分的恒等式,将多个分数(或此处的t分布)乘积表示为单纯形上的积分。利用此单纯形潜空间,可从乘积专家中抽取加权样本,供BBVI用于寻找最接近目标密度的泊松专家组合。针对一组专家,我们推导出一种迭代优化算法,以确定其在乘积中的几何权重。每轮迭代通过最小化正则化的费希尔散度,在当前近似样本上匹配变分密度与目标密度的得分。该最小化问题可转化为凸二次规划,并在一般条件下证明更新过程以指数速度收敛至近优权重。我们在多种合成与真实世界目标分布上验证了该方法的有效性。

原文摘要 · Abstract (English)

We introduce a highly expressive yet distinctly tractable family for black-box variational inference (BBVI). Each member of this family is a weighted product of experts (PoE), and each weighted expert in the product is proportional to a multivariate $t$-distribution. These products of experts can model distributions with skew, heavy tails, and multiple modes, but to use them for BBVI, we must be able to sample from their densities. We show how to do this by reformulating these products of experts as latent variable models with auxiliary Dirichlet random variables. These Dirichlet variables emerge from a Feynman identity, originally developed for loop integrals in quantum field theory, that expresses the product of multiple fractions (or in our case, $t$-distributions) as an integral over the simplex. We leverage this simplicial latent space to draw weighted samples from these products of experts -- samples which BBVI then uses to find the PoE that best approximates a target density. Given a collection of experts, we derive an iterative procedure to optimize the exponents that determine their geometric weighting in the PoE. At each iteration, this procedure minimizes a regularized Fisher divergence to match the scores of the variational and target densities at a batch of samples drawn from the current approximation. This minimization reduces to a convex quadratic program, and we prove under general conditions that these updates converge exponentially fast to a near-optimal weighting of experts. We conclude by evaluating this approach on a variety of synthetic and real-world target distributions.

变分推断乘积专家t分布得分匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。