arXiv:2601.22441stat.MLcs.LG2026-01被引 1

用学习到的统计量改进贝叶斯推断,解决模拟数据难以计算似然的问题。

Simulation-based Bayesian inference with ameliorative learned summary statistics -- Part I

  • 通过变换技术融合观测与模拟数据的统计量,提升推断效率。
  • 在弱相关数据下仍保持良好推断性能,适用于复杂模型。
  • 可分布式实现,适合大规模数据和复杂模拟场景。

本文为两篇系列论文的第一部分,研究基于模拟的贝叶斯推断框架,其中学习得到的总结统计量作为具有修正效果的经验似然,在真实似然函数无法显式表达或计算不可行时发挥作用。特别地,利用受矩约束的Cressie-Read分歧准则进行转换,以压缩观测数据与模拟输出之间的学习统计量,同时保持推断的统计效能。该转换使模拟输出能基于观测数据条件化,从而在特定样本子集上执行推断,这些子集被视为具有经验相关性或重要性。此外,该框架可进一步扩展至处理弱依赖观测数据。最后,该方法适合分布式计算实现,数据到学习统计量的转换及贝叶斯推断任务可统一为分布式推断问题,借助分布式优化与MCMC算法支持大规模数据与复杂模拟模型。

原文摘要 · Abstract (English)

This paper, which is Part 1 of a two-part paper series, considers a simulation-based inference with learned summary statistics, in which such a learned summary statistic serves as an empirical-likelihood with ameliorative effects in the Bayesian setting, when the exact likelihood function associated with the observation data and the simulation model is difficult to obtain in a closed form or computationally intractable. In particular, a transformation technique which leverages the Cressie-Read discrepancy criterion under moment restrictions is used for summarizing the learned statistics between the observation data and the simulation outputs, while preserving the statistical power of the inference. Here, such a transformation of data-to-learned summary statistics also allows the simulation outputs to be conditioned on the observation data, so that the inference task can be performed over certain sample sets of the observation data that are considered as an empirical relevance or believed to be particular importance. Moreover, the simulation-based inference framework discussed in this paper can be extended further, and thus handling weakly dependent observation data. Finally, we remark that such an inference framework is suitable for implementation in distributed computing, i.e., computational tasks involving both the data-to-learned summary statistics and the Bayesian inferencing problem can be posed as a unified distributed inference problem that will exploit distributed optimization and MCMC algorithms for supporting large datasets associated with complex simulation models.

贝叶斯推断模拟推断分布式计算学习统计量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。