证明了贝叶斯神经网络在特定条件下可同时实现最优与可接受性。
Minimaxity and Admissibility of Bayesian Neural Networks
- 通过在先验方差上引入超先验,构造出满足极小极大性的决策规则。
- 固定先验尺度下,原贝叶斯神经网络不满足极小极大性,而新方法可达到。
- 理论结果适用于均方误差和预测密度估计,适合关注模型最优性的研究者。
贝叶斯神经网络(BNNs)为深度学习中的推断提供了自然的概率框架。尽管广泛应用,其最优性在统计决策理论视角下仍缺乏深入研究。本文研究了在二次损失下的正态位置模型中,由深层全连接ReLU BNNs诱导的决策规则。我们发现:在固定先验尺度下,对应的贝叶斯决策规则并非极小极大。随后,我们对BNN先验的有效输出方差引入超先验,得到一个超调和的平方根边际密度,从而确保所导出的决策规则同时具有可接受性和极小极大性。进一步将结果从二次损失推广至以Kullback-Leibler损失衡量的预测密度估计问题。最后,通过数值模拟验证了理论结论。
原文摘要 · Abstract (English)
Bayesian neural networks (BNNs) offer a natural probabilistic formulation for inference in deep learning models. Despite their popularity, their optimality has received limited attention through the lens of statistical decision theory. In this paper, we study decision rules induced by deep, fully connected feedforward ReLU BNNs in the normal location model under quadratic loss. We show that, for fixed prior scales, the induced Bayes decision rule is not minimax. We then propose a hyperprior on the effective output variance of the BNN prior that yields a superharmonic square-root marginal density, establishing that the resulting decision rule is simultaneously admissible and minimax. We further extend these results from the quadratic loss setting to the predictive density estimation problem with Kullback--Leibler loss. Finally, we validate our theoretical findings numerically through simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。