优化计算顺序与动态分解,大幅提升高斯过程模型的采样效率。
Fast Riemannian-manifold Hamiltonian Monte Carlo for hierarchical Gaussian-process models
- 根据模型结构优化计算顺序,动态编程特征分解
- 在模拟数据和真实医疗支出数据上实现有效后验采样
- 适合需要高效贝叶斯推断的复杂非线性建模者
基于高斯过程的分层贝叶斯模型可用于描述真实世界数据中变量间的复杂非线性依赖关系。然而,针对这些模型的有效蒙特卡洛推断算法尚未建立,仅少数简单情况例外。本研究表明,通过根据模型结构优化计算顺序并动态编程特征分解,相比现有程序库的缓慢推断,黎曼流形哈密顿蒙特卡洛(RMHMC)性能可显著提升。这一改进无法通过依赖朴素自动微分的现有库实现。数值实验显示,RMHMC在模拟数据的贝叶斯逻辑回归及美国全国医疗支出数据的倾向函数估计中,能有效从后验分布采样,计算模型证据。结果为使用高斯过程分析真实世界数据提供了有效的蒙特卡洛算法基础,并强调需开发可定制库,支持动态编程对象及根据模型结构精细优化自动微分方式。
原文摘要 · Abstract (English)
Hierarchical Bayesian models based on Gaussian processes are considered useful for describing complex nonlinear statistical dependencies among variables in real-world data. However, effective Monte Carlo algorithms for inference with these models have not yet been established, except for several simple cases. In this study, we show that, compared with the slow inference achieved with existing program libraries, the performance of Riemannian-manifold Hamiltonian Monte Carlo (RMHMC) can be drastically improved by optimising the computation order according to the model structure and dynamically programming the eigendecomposition. This improvement cannot be achieved when using an existing library based on a naive automatic differentiator. We numerically demonstrate that RMHMC effectively samples from the posterior, allowing the calculation of model evidence, in a Bayesian logistic regression on simulated data and in the estimation of propensity functions for the American national medical expenditure data using several Bayesian multiple-kernel models. These results lay a foundation for implementing effective Monte Carlo algorithms for analysing real-world data with Gaussian processes, and highlight the need to develop a customisable library set that allows users to incorporate dynamically programmed objects and finely optimises the mode of automatic differentiation depending on the model structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。