arXiv:2409.01574stat.COcs.LG2024-09中稿 · ICML被引 3

用强化学习动态调温度,让马尔可夫链采样更快收敛

Policy Gradients for Optimal Parallel Tempering MCMC

  • 用策略梯度算法自动调整多链温度
  • 在基准测试中自相关时间更短
  • 适合需要高效采样的复杂分布场景

并行退火是一种马尔可夫链蒙特卡洛的元算法,通过多个链采样目标分布的退火版本,提升多模态分布的混合效率。其效果高度依赖链温度的选择。本文提出一种基于策略梯度的自适应温度选择算法,在采样过程中动态调整温度。实验表明,该方法在基准分布上相比传统的几何间隔温度和均匀接受率方案,能实现更低的积分自相关时间。

原文摘要 · Abstract (English)

Parallel tempering is a meta-algorithm for Markov Chain Monte Carlo that uses multiple chains to sample from tempered versions of the target distribution, enhancing mixing in multi-modal distributions that are challenging for traditional methods. The effectiveness of parallel tempering is heavily influenced by the selection of chain temperatures. Here, we present an adaptive temperature selection algorithm that dynamically adjusts temperatures during sampling using a policy gradient approach. Experiments demonstrate that our method can achieve lower integrated autocorrelation times compared to traditional geometrically spaced temperatures and uniform acceptance rate schemes on benchmark distributions.

MCMC采样优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。