arXiv:2412.00545stat.MLcs.LG2024-12

为离散分布的粒子近似找到最优加权方案,显著提升各类粒子方法性能。

Optimal Particle-based Approximation of Discrete Distributions (OPAD)

  • 基于已有粒子计算值,推导出最小化KL散度的唯一最优权重
  • 在变量选择与结构学习任务中,重加权后近似精度持续且显著提升
  • 无需额外计算成本,可通用改进现有粒子方法,适合概率建模研究者

粒子方法(如马尔可夫链蒙特卡洛和序贯蒙特卡洛)通过一组带权粒子逼近目标概率分布。本文证明:当目标分布为离散时,对任意粒子集合,存在唯一加权方式使粒子近似与目标分布的KL散度最小,任何其他加权机制(如基于马尔可夫链重复次数的MCMC加权)均次优。该结论不依赖目标分布形式或粒子生成过程,仅要求目标为离散分布。最优权重可通过现有方法已计算的值确定,仅需微小修改即可实现性能提升,无额外计算开销。实验在贝叶斯变量选择与贝叶斯结构学习等典型离散分布应用中验证,结果表明重加权能一致且显著改善粒子近似效果。

原文摘要 · Abstract (English)

Particle-based methods include a variety of techniques, such as Markov Chain Monte Carlo (MCMC) and Sequential Monte Carlo (SMC), for approximating a probabilistic target distribution with a set of weighted particles. In this paper, we prove that for any set of particles, there is a unique weighting mechanism that minimizes the Kullback-Leibler (KL) divergence of the (particle-based) approximation from the target distribution, when that distribution is discrete -- any other weighting mechanism (e.g. MCMC weighting that is based on particles' repetitions in the Markov chain) is sub-optimal with respect to this divergence measure. Our proof does not require any restrictions either on the target distribution, or the process by which the particles are generated, other than the discreteness of the target. We show that the optimal weights can be determined based on values that any existing particle-based method already computes; As such, with minimal modifications and no extra computational costs, the performance of any particle-based method can be improved. Our empirical evaluations are carried out on important applications of discrete distributions including Bayesian Variable Selection and Bayesian Structure Learning. The results illustrate that our proposed reweighting of the particles improves any particle-based approximation to the target distribution consistently and often substantially.

粒子方法概率近似贝叶斯推断优化加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。