arXiv:2412.06082cs.CV2024-12被引 10

验证视觉大模型能否可靠做不确定性预测。

Are foundation models for computer vision good conformal predictors?

  • 用分位数校准法提升大模型置信度反而降低预测效率。
  • 视觉-语言模型少样本微调后,预测置信度更高。
  • 自适应分位数方法在复杂场景下仍保持覆盖保证。

自监督与对比学习的进展使视觉基础模型在各类任务中达到前所未有的性能水平,正广泛应用于风险敏感的高价值场景。然而,确保其安全部署需深入理解其不确定性建模能力,该问题尚未受到足够关注。本文系统研究了视觉与视觉-语言基础模型在分位数预测(Conformal Prediction, CP)框架下的表现。在多个主流图像分类基准、典型基础模型及三种CP方法上实验发现,基础模型尤其适合基于视觉变换器的校准流程。此外,校准模型置信度虽常用于改进不确定性量化,却会降低自适应分位数方法的预测集效率。少样本适配视觉-语言模型可显著提升其分位数得分,优于零样本预测。最后,我们的实证研究显示,自适应分位数方法(APS)在多重挑战性但真实的场景中均能保持边际覆盖保证,展现出巨大潜力。

原文摘要 · Abstract (English)

Recent advances in self-supervision and contrastive learning have brought the performance of foundation models to unprecedented levels in a variety of tasks. Fueled by this progress, these models are becoming the prevailing approach for a wide array of real-world vision problems, including risk-sensitive and high-stakes applications. However, ensuring safe deployment in these scenarios requires a more comprehensive understanding of their uncertainty modeling capabilities, which has received little attention. In this work, we delve into the behaviour of vision and vision-language foundation models under Conformal Prediction (CP), a statistical framework that provides theoretical guarantees of marginal coverage of the true class. Across extensive experiments including popular vision classification benchmarks, well-known foundation vision models, and three CP methods, our findings reveal that foundation models are well-suited for conformalization procedures, particularly those integrating Vision Transformers. We also show that calibrating the confidence predictions of these models, a popular strategy to improve their uncertainty quantification, actually leads to efficiency degradation of the conformal set on adaptive CP methods. Furthermore, few-shot adaptation of Vision-Language Models (VLMs) to downstream tasks, whose popularity is surging, enhances conformal scores compared to zero-shot predictions. Last, our empirical study exposes APS as particularly promising in the context of vision foundation models, as it does not violate the marginal coverage guarantees across multiple challenging, yet realistic scenarios.

视觉大模型不确定性分位数预测稳健性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。