用奖励机制引导粒子采样,高效逼近未知密度分布
R-ParVI: Particle-based variational inference through lens of rewards
- 以奖励驱动粒子流,无需梯度即可采样部分已知密度
- 粒子通过随机扰动与奖励机制,避免聚集并保持多样性
- 适合贝叶斯推断和生成建模,速度快且可扩展
提出一种奖励引导的无梯度粒子变分推断方法 R-ParVI,用于采样部分已知密度(如仅知常数倍)。R-ParVI 将采样问题建模为由奖励驱动的粒子流动:粒子从先验分布出发,在参数空间中移动,其路径由融合目标密度评估的奖励机制决定,稳态下的粒子配置逼近目标分布几何结构。通过随机扰动与奖励机制模拟粒子与环境的交互,使粒子向高密度区域迁移的同时保持多样性(如防止塌陷成簇)。R-ParVI 为贝叶斯推断与生成建模等概率模型提供快速、灵活、可扩展且随机的采样与推断方案。
原文摘要 · Abstract (English)
A reward-guided, gradient-free ParVI method, \textit{R-ParVI}, is proposed for sampling partially known densities (e.g. up to a constant). R-ParVI formulates the sampling problem as particle flow driven by rewards: particles are drawn from a prior distribution, navigate through parameter space with movements determined by a reward mechanism blending assessments from the target density, with the steady state particle configuration approximating the target geometry. Particle-environment interactions are simulated by stochastic perturbations and the reward mechanism, which drive particles towards high density regions while maintaining diversity (e.g. preventing from collapsing into clusters). R-ParVI offers fast, flexible, scalable and stochastic sampling and inference for a class of probabilistic models such as those encountered in Bayesian inference and generative modelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。