arXiv:2505.24693cs.CV2025-05CVPR被引 13

用新方法提升零样本模型的预测可靠性,让结果更可信且更快。

Conformal Prediction for Zero-Shot Models

  • 基于最优传输思想,融合校准与查询数据,缓解预训练与任务间的领域差异
  • 在15个数据集上使预测集合效率提升最高达20%,速度比主流方法快15倍
  • 适合关注模型可靠性、需快速生成置信预测的零样本应用开发者

大规模预训练的视觉-语言模型展现出前所未有的适应性和泛化能力。尽管其判别性能被广泛研究,但其可靠性和不确定性仍被忽视。本文在分片共形预测框架下,探究CLIP模型的能力,该框架基于少量标注校准集为黑箱模型提供理论保证。与传统视觉分类器的共形预测研究不同,基础模型具有独特特征:其仅在一次不可见源域上进行预训练,与目标任务存在领域漂移,这会降低共形集合的效率并带来额外挑战。为此,我们提出Conf-OT,一种在联合校准与查询集上进行归纳学习的迁移学习设定。通过求解最优传输问题,该方法在不需额外数据划分的前提下弥合了预训练与适应阶段之间的领域差距,同时保持覆盖率保证。我们在15个数据集和三种非一致性度量上全面验证该策略,Conf-OT在集合效率上实现最高达20%的相对提升,且速度比主流归纳方法快15倍。

原文摘要 · Abstract (English)

Vision-language models pre-trained at large scale have shown unprecedented adaptability and generalization to downstream tasks. Although its discriminative potential has been widely explored, its reliability and uncertainty are still overlooked. In this work, we investigate the capabilities of CLIP models under the split conformal prediction paradigm, which provides theoretical guarantees to black-box models based on a small, labeled calibration set. In contrast to the main body of literature on conformal predictors in vision classifiers, foundation models exhibit a particular characteristic: they are pre-trained on a one-time basis on an inaccessible source domain, different from the transferred task. This domain drift negatively affects the efficiency of the conformal sets and poses additional challenges. To alleviate this issue, we propose Conf-OT, a transfer learning setting that operates transductive over the combined calibration and query sets. Solving an optimal transport problem, the proposed method bridges the domain gap between pre-training and adaptation without requiring additional data splits but still maintaining coverage guarantees. We comprehensively explore this conformal prediction strategy on a broad span of 15 datasets and three non-conformity scores. Conf-OT provides consistent relative improvements of up to 20% on set efficiency while being 15 times faster than popular transductive approaches.

共形预测零样本可靠性CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。