深度集成的贝叶斯神经网络在分布内数据上反而表现更差。
On Local Posterior Structure in Deep Ensembles
- 对比贝叶斯神经网络与深度集成,研究发现大模型集成反而更优。
- 当集成规模增大时,深度集成在分布内任务上持续优于贝叶斯集成。
- 适合关注不确定性量化与模型校准的研究者参考。
贝叶斯神经网络(BNN)相比点估计方法(如最大后验估计)能提升模型校准与预测不确定性量化。类似地,深度集成(DE)也改善了校准效果,因此自然推测贝叶斯神经网络的深度集成(DE-BNN)应带来更大改进。本文在多个数据集、神经网络架构及BNN近似方法上系统研究该假设,出人意料地发现:当集成规模足够大时,深度集成(DE)在分布内数据上的表现始终优于DE-BNN。为揭示此现象,我们进行了多项敏感性与消融实验。此外,我们证明尽管DE-BNN在分布外指标上优于DE,但代价是分布内性能下降。作为最终贡献,我们开源了大量训练好的模型,以促进该方向的进一步研究。
原文摘要 · Abstract (English)
Bayesian Neural Networks (BNNs) often improve model calibration and predictive uncertainty quantification compared to point estimators such as maximum-a-posteriori (MAP). Similarly, deep ensembles (DEs) are also known to improve calibration, and therefore, it is natural to hypothesize that deep ensembles of BNNs (DE-BNNs) should provide even further improvements. In this work, we systematically investigate this across a number of datasets, neural network architectures, and BNN approximation methods and surprisingly find that when the ensembles grow large enough, DEs consistently outperform DE-BNNs on in-distribution data. To shine light on this observation, we conduct several sensitivity and ablation studies. Moreover, we show that even though DE-BNNs outperform DEs on out-of-distribution metrics, this comes at the cost of decreased in-distribution performance. As a final contribution, we open-source the large pool of trained models to facilitate further research on this topic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。