贝叶斯神经网络可实现自信预测,揭示了不确定性量化的新机制。
Can Bayesian Neural Networks Make Confident Predictions?
- 通过离散化权重先验,精确刻画后验预测为高斯混合分布。
- 发现低训练误差下存在多个后验模式,导致预测多模态现象。
- 适用于研究模型不确定性、训练规模与网络结构关系的科研人员。
贝叶斯推断为神经网络预测提供了严谨的不确定性量化框架,但其应用受限于参数后验分布难以完全表征以及后验预测分布解释性差。本文证明,在对隐层权重采用离散化先验时,可将后验预测分布精确表示为高斯混合形式。该设定下,可定义产生相同似然(训练误差)的参数值等价类,并将其与网络缩放范式(由训练样本量、各层规模及输出层参数数量之比定义)关联起来。特别地,识别出在低训练误差下仍对应后验预测分布多个模式的参数配置,揭示了非单峰后验近似可能存在的偏差。同时,通过评估不同缩放范式下后验预测的收缩情况,刻画了模型从数据中学习的能力。
原文摘要 · Abstract (English)
Bayesian inference promises a framework for principled uncertainty quantification of neural network predictions. Barriers to adoption include the difficulty of fully characterizing posterior distributions on network parameters and the interpretability of posterior predictive distributions. We demonstrate that under a discretized prior for the inner layer weights, we can exactly characterize the posterior predictive distribution as a Gaussian mixture. This setting allows us to define equivalence classes of network parameter values which produce the same likelihood (training error) and to relate the elements of these classes to the network's scaling regime -- defined via ratios of the training sample size, the size of each layer, and the number of final layer parameters. Of particular interest are distinct parameter realizations that map to low training error and yet correspond to distinct modes in the posterior predictive distribution. We identify settings that exhibit such predictive multimodality, and thus provide insight into the accuracy of unimodal posterior approximations. We also characterize the capacity of a model to "learn from data" by evaluating contraction of the posterior predictive in different scaling regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。