用E值设计停止规则,让贝叶斯深度集成采样更省时高效。
Towards E-Value Based Stopping Rules for Bayesian Deep Ensembles
- 基于E值构建序贯检验,动态判断采样是否继续。
- 实验表明只需全预算的一小部分即可达到显著提升。
- 适合需要高效不确定性量化且资源受限的场景。
贝叶斯深度集成(BDEs)结合了深度集成(DEs)的鲁棒性与多链MCMC的灵活性,是深度学习中不确定性量化的重要方法。尽管DEs在多数场景下成本可控,但贝叶斯神经网络的长时采样仍可能代价过高。已有研究表明,在优化后的DE基础上追加采样可带来显著性能提升。这引出一个关键问题:采样应持续多久才能获得明显改进?为此,本文提出一种基于E值的停止规则。将集成构建建模为序贯的任意时间有效性假设检验,提供了一种严谨方法来判断是否拒绝“MCMC未优于强基线”这一零假设,从而实现早期停止。我们在多种设置下验证该方法,结果表明其有效性,并揭示通常仅需极少部分完整采样预算即可取得显著增益。
原文摘要 · Abstract (English)
Bayesian Deep Ensembles (BDEs) represent a powerful approach for uncertainty quantification in deep learning, combining the robustness of Deep Ensembles (DEs) with flexible multi-chain MCMC. While DEs are affordable in most deep learning settings, (long) sampling of Bayesian neural networks can be prohibitively costly. Yet, adding sampling after optimizing the DEs has been shown to yield significant improvements. This leaves a critical practical question: How long should the sequential sampling process continue to yield significant improvements over the initial optimized DE baseline? To tackle this question, we propose a stopping rule based on E-values. We formulate the ensemble construction as a sequential anytime-valid hypothesis test, providing a principled way to decide whether or not to reject the null hypothesis that MCMC offers no improvement over a strong baseline, to early stop the sampling. Empirically, we study this approach for diverse settings. Our results demonstrate the efficacy of our approach and reveal that only a fraction of the full-chain budget is often required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。