arXiv:2506.09593cs.LG2025-06被引 5

基础模型预测更可靠,但需注意校准方法在分布外时可能失效

Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks

  • 发现基础模型在分布内预测偏保守,导致校准误差更高
  • 在分布外情况下,校准性能反而优于传统模型
  • 后处理校准法在分布内有效,但在严重分布偏移下可能适得其反

可靠的不确定性校准对于高风险场景中部署深度神经网络至关重要。深度神经网络在分布偏移下普遍存在系统性过自信问题。尽管像ConvNeXt、EVA和BEiT这样的基础模型在预测性能上显著提升,其校准特性仍缺乏研究。本文全面分析了基础模型的校准行为,发现这些模型在分布内预测倾向于低估置信度,导致更高的校准误差,但在分布外情况下表现出更好的校准能力。此外,我们证明基础模型对后处理校准技术在分布内高度敏感,可有效缓解低估偏差;然而,在严重分布偏移下,这些方法逐渐不可靠,甚至可能产生反效果。研究揭示了架构与训练创新对校准的复杂非单调影响,挑战了持续改进的既有认知。

原文摘要 · Abstract (English)

Reliable uncertainty calibration is essential for safely deploying deep neural networks in high-stakes applications. Deep neural networks are known to exhibit systematic overconfidence, especially under distribution shifts. Although foundation models such as ConvNeXt, EVA and BEiT have demonstrated significant improvements in predictive performance, their calibration properties remain underexplored. This paper presents a comprehensive investigation into the calibration behavior of foundation models, revealing insights that challenge established paradigms. Our empirical analysis shows that these models tend to be underconfident in in-distribution predictions, resulting in higher calibration errors, while demonstrating improved calibration under distribution shifts. Furthermore, we demonstrate that foundation models are highly responsive to post-hoc calibration techniques in the in-distribution setting, enabling practitioners to effectively mitigate underconfidence bias. However, these methods become progressively less reliable under severe distribution shifts and can occasionally produce counterproductive results. Our findings highlight the complex, non-monotonic effects of architectural and training innovations on calibration, challenging established narratives of continuous improvement.

模型校准基础模型分布外不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。