针对计算昂贵且不规则的似然函数,提出高效精确的采样新方法。
Efficient MCMC Sampling with Expensive-to-Compute and Irregular Likelihoods
- 用子集评估降低计算开销,结合数据驱动代理替代泰勒展开。
- 在固定计算预算下,新方法采样误差最低,优于现有算法。
- 适合高维复杂似然场景,如疾病建模与不规则分布采样任务。
当贝叶斯推断中的似然函数不规则且计算成本高昂时,马尔可夫链蒙特卡洛(MCMC)采样面临挑战。本文探索了利用子集评估减少计算开销的采样算法,并针对梯度不可靠或不可用的情况,改进了子集采样器。为此,引入数据驱动代理替代泰勒展开,并设计新型计算成本感知自适应控制器。在具有挑战性的疾病建模任务及具有类似似然曲面不规则性的可配置任务上进行了广泛评估。结果表明,改进版层级重要性与嵌套训练样本法(HINTS)结合自适应提议与数据驱动代理,在固定计算预算下取得了最低的采样误差。研究结论表明,子集评估能提供廉价且自然调节的探索能力,数据驱动代理可在已探索状态空间区域有效预筛提议,二者通过层级延迟接受机制实现高效且精确的采样。
原文摘要 · Abstract (English)
Bayesian inference with Markov Chain Monte Carlo (MCMC) is challenging when the likelihood function is irregular and expensive to compute. We explore several sampling algorithms that make use of subset evaluations to reduce computational overhead. We adapt the subset samplers for this setting where gradient information is not available or is unreliable. To achieve this, we introduce data-driven proxies in place of Taylor expansions and define a novel computation-cost aware adaptive controller. We undertake an extensive evaluation for a challenging disease modelling task and a configurable task with similar irregularity in the likelihood surface. We find our improved version of Hierarchical Importance with Nested Training Samples (HINTS), with adaptive proposals and a data-driven proxy, obtains the best sampling error in a fixed computational budget. We conclude that subset evaluations can provide cheap and naturally-tempered exploration, while a data-driven proxy can pre-screen proposals successfully in explored regions of the state space. These two elements combine through hierarchical delayed acceptance to achieve efficient, exact sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。