arXiv:2601.07944stat.MLcs.LG2026-01

用神经网络加速贝叶斯推断,训练一次,处处可用。

Neural Architectures for Amortized Bayesian Inference: Statistical Foundations and Empirical Assessments

  • 用前馈网络、Deep Sets、Transformer等模型实现可复用的贝叶斯推断
  • 在不同数据规模、噪声分布、稀疏性下保持稳定误差和可靠不确定性估计
  • 适合需要快速推理且需量化不确定性的机器学习应用

自本世纪初以来,近似贝叶斯推断随着新计算技术的引入,逐步发展以应对日益复杂的大规模预测问题。深度神经网络与基础模型的成功催生了一种新范式:通过大规模学习的预测器实现贝叶斯推断的摊销。在摊销推断中,前期需大量计算训练神经网络,但后续可在多种任务上以极低开销生成近似后验或预测。相比传统贝叶斯方法因每次新数据集需重复似然计算与蒙特卡洛步骤而计算昂贵,摊销推断显著降低部署成本。尽管摊销推断日益流行,其统计意义及在贝叶斯框架中的定位仍不清晰。本文从统计视角分析前馈网络、Deep Sets、Transformer等主流神经架构如何自然支持摊销贝叶斯推断,探讨其结构化近似与概率推理能力如何在广泛部署场景中实现可控泛化误差,并揭示其在贝叶斯计算中的潜力。通过模拟实验,在不同样本量、噪声分布族、稀疏性水平和多模态条件下评估摊销推断的精度、鲁棒性与不确定性量化表现,明确其优势与局限。

原文摘要 · Abstract (English)

Since the turn of the century, approximate Bayesian inference has steadily evolved as new computational techniques have been incorporated to handle increasingly complex, large-scale predictive problems. The recent success of deep neural networks and foundation models has now given rise to a new paradigm in statistical modeling, in which Bayesian inference can be amortized through large-scale learned predictors. In amortized inference, substantial computation is required at the beginning to train a neural network, but it can subsequently produce approximate posteriors or predictions at much lower computational cost across a wide range of tasks. While the typical Bayesian inference procedures are computationally expensive due to repeated likelihood calculations and Monte Carlo steps for each new dataset, amortized inference provides a much lower computational cost at deployment. Despite the growing popularity of amortized inference, its statistical interpretation and position within Bayesian inference remain poorly explored. In this paper, we present a statistical perspective on several major neural architectures, including feedforward networks, Deep Sets, and Transformers, and examine how they naturally support amortized Bayesian inference. We explore how these models perform structured approximation and also probabilistic reasoning in ways that yield controlled generalization error throughout a wide range of deployment scenarios, and how these properties can be harnessed for Bayesian computation. Via simulation studies, we evaluate the accuracy, robustness, and uncertainty quantification of amortized inference across varying sample sizes, varying noise distributional families, varying sparsity levels, and multimodality, highlighting its strengths and limitations.

贝叶斯推断神经网络摊销推理不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。