将浅层贝叶斯神经网络与高斯过程统一,实现稳定推断与高效计算。
From Shallow Bayesian Neural Networks to Gaussian Processes: General Convergence, Identifiability and Scalable Inference
- 通过松弛假设建立BNN到GP的通用收敛理论
- 提出凸混合核函数,具备严格可识别性与正定性
- 基于Nyström近似实现低成本高精度训练与预测
本文研究浅层贝叶斯神经网络(BNN)在极限情况下的行为,揭示其与高斯过程(GP)之间的联系,重点关注统计建模、参数可识别性与可扩展推断。我们放松了以往理论中的假设,建立了从BNN到GP的通用收敛结果,并比较了不同参数化下极限GP模型的差异。在此基础上,提出一种新的协方差函数,由四种常见激活函数诱导的成分构成凸混合形式,刻画了其正定性以及在不同输入设计下的严格与实用可识别性。针对计算问题,开发了一种基于Nyström近似的可扩展最大后验(MAP)训练与预测方法,证明了Nyström秩和锚点选择可调节计算成本与精度的权衡。在受控模拟与真实世界表格数据集上的实验表明,该方法能获得稳定的超参数估计,并在合理计算成本下达到有竞争力的预测性能。
原文摘要 · Abstract (English)
In this work, we study scaling limits of shallow Bayesian neural networks (BNNs) via their connection to Gaussian processes (GPs), with an emphasis on statistical modeling, identifiability, and scalable inference. We first establish a general convergence result from BNNs to GPs by relaxing assumptions used in prior formulations, and we compare alternative parameterizations of the limiting GP model. Building on this theory, we propose a new covariance function defined as a convex mixture of components induced by four widely used activation functions, and we characterize key properties including positive definiteness and both strict and practical identifiability under different input designs. For computation, we develop a scalable maximum a posterior (MAP) training and prediction procedure using a Nyström approximation, and we show how the Nyström rank and anchor selection control the cost-accuracy trade-off. Experiments on controlled simulations and real-world tabular datasets demonstrate stable hyperparameter estimates and competitive predictive performance at realistic computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。