arXiv:2409.02628cs.LGstat.ML2024-09被引 15

大模型越复杂,不确定性反而越差,研究发现可通过隐式集成恢复可靠性。

(Implicit) Ensembles of Ensembles: Epistemic Uncertainty Collapse in Large Models

  • 提出隐式集成机制解释大模型不确定性衰减现象。
  • 实验显示从MLP到ViT,模型越宽不确定性越弱。
  • 开发隐式集成提取技术,可从大模型中还原多样子模型。

认知不确定性对安全关键应用和数据采集任务至关重要。然而我们发现深度学习模型存在一个关键现象:随着模型复杂度增加,认知不确定性出现坍塌,挑战了大模型必然提供更好不确定性量化这一假设。本文提出隐式集成作为可能的解释。通过理论分析与实验,验证了显式集成组合中的不确定性坍塌,并在多种架构(从简单MLP到先进视觉模型如ResNet和Vision Transformer)中发现类似现象。进一步开发隐式集成提取技术,将大模型分解为多样化子模型,成功恢复认知不确定性。这些发现对不确定性估计具有重要启示。

原文摘要 · Abstract (English)

Epistemic uncertainty is crucial for safety-critical applications and data acquisition tasks. Yet, we find an important phenomenon in deep learning models: an epistemic uncertainty collapse as model complexity increases, challenging the assumption that larger models invariably offer better uncertainty quantification. We introduce implicit ensembling as a possible explanation for this phenomenon. To investigate this hypothesis, we provide theoretical analysis and experiments that demonstrate uncertainty collapse in explicit ensembles of ensembles and show experimental evidence of similar collapse in wider models across various architectures, from simple MLPs to state-of-the-art vision models including ResNets and Vision Transformers. We further develop implicit ensemble extraction techniques to decompose larger models into diverse sub-models, showing we can thus recover epistemic uncertainty. We explore the implications of these findings for uncertainty estimation.

不确定性大模型隐式集成深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。