深度集成其实暗中实现了经验贝叶斯,用数据学出先验。
Deep Ensembles Secretly Perform Empirical Bayes
- 用数据隐式学习先验,实现精确贝叶斯平均
- 揭示了集成模型性能优于贝叶斯神经网络的理论原因
- 适合关心模型不确定性与集成原理的研究者
量化神经网络的不确定性是众多应用中的关键问题。当前主流方法包括贝叶斯神经网络(BNNs)和深度集成(deep ensembles)。尽管两者有相似之处,通常被认为无正式关联且本质不同:BNNs因基于贝叶斯范式被视为更严谨,而集成则被视作更随意;然而,深度集成在实践中常优于BNNs,却缺乏合理解释。本文揭示:深度集成实际上执行了精确的贝叶斯平均,其后验由一个隐式学习的数据依赖先验得出。换言之,深度集成本质上是经验贝叶斯方法,先验由数据学习。该视角带来两大好处:(i) 理论上解释了深度集成的优异表现;(ii) 分析学习到的先验发现其为点质量混合——这种强先验有助于理解集成中的观测现象。本工作重新定义了对深度集成的理解,不仅具有理论价值,也可能催生未来模型改进。
原文摘要 · Abstract (English)
Quantifying uncertainty in neural networks is a highly relevant problem which is essential to many applications. The two predominant paradigms to tackle this task are Bayesian neural networks (BNNs) and deep ensembles. Despite some similarities between these two approaches, they are typically surmised to lack a formal connection and are thus understood as fundamentally different. BNNs are often touted as more principled due to their reliance on the Bayesian paradigm, whereas ensembles are perceived as more ad-hoc; yet, deep ensembles tend to empirically outperform BNNs, with no satisfying explanation as to why this is the case. In this work we bridge this gap by showing that deep ensembles perform exact Bayesian averaging with a posterior obtained with an implicitly learned data-dependent prior. In other words deep ensembles are Bayesian, or more specifically, they implement an empirical Bayes procedure wherein the prior is learned from the data. This perspective offers two main benefits: (i) it theoretically justifies deep ensembles and thus provides an explanation for their strong empirical performance; and (ii) inspection of the learned prior reveals it is given by a mixture of point masses -- the use of such a strong prior helps elucidate observed phenomena about ensembles. Overall, our work delivers a newfound understanding of deep ensembles which is not only of interest in it of itself, but which is also likely to generate future insights that drive empirical improvements for these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。