arXiv:2601.16936cs.LG2026-01中稿 · the 1st workshop o…被引 4

BatchEnsemble看似高效,实则接近单模型,不确定性估计不可靠。

Is BatchEnsemble a Single Model? On Calibration and Diversity of Efficient Ensembles

  • 用低秩扰动共享主网络,模拟集成学习
  • 在多个数据集上表现与单模型无异
  • 适合对计算资源敏感但不重精度的场景

在资源受限和低延迟场景中,需高效获取不确定性估计。深度集成虽能提供可靠的认知不确定性(EU),但需训练多个完整模型。BatchEnsemble通过在共享基础网络上施加可学习的低秩扰动,以极低参数量和内存开销实现类集成的EU。我们发现,BatchEnsemble不仅性能弱于深度集成,在CIFAR10/10C/SVHN上的准确率、校准度及分布外(OOD)检测表现均与单模型基线相近。在MNIST上的受控实验表明,各成员在函数空间和参数空间中几乎完全一致,表明其难以实现多样化的预测模式。因此,BatchEnsemble的行为更像单模型而非真正集成。

原文摘要 · Abstract (English)

In resource-constrained and low-latency settings, uncertainty estimates must be efficiently obtained. Deep Ensembles provide robust epistemic uncertainty (EU) but require training multiple full-size models. BatchEnsemble aims to deliver ensemble-like EU at far lower parameter and memory cost by applying learned rank-1 perturbations to a shared base network. We show that BatchEnsemble not only underperforms Deep Ensembles but closely tracks a single model baseline in terms of accuracy, calibration and out-of-distribution (OOD) detection on CIFAR10/10C/SVHN. A controlled study on MNIST finds members are near-identical in function and parameter space, indicating limited capacity to realize distinct predictive modes. Thus, BatchEnsemble behaves more like a single model than a true ensemble.

集成学习不确定性估计模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。