提出新方法精准识别CLIP模型依赖的图像区域,提升解释力与鲁棒性。
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- 用视觉块聚类分析模型关注点,量化语义区域重要性。
- 在删除-AUC指标上超越旧方法两倍以上,达新基准水平。
- 可区分前景与背景依赖,适合研究模型偏差与可靠性者使用。
对比视觉语言模型(如CLIP)虽具备强零样本识别能力,但仍易受无关背景干扰。本文提出基于聚类的概念重要性(CCI)方法,利用CLIP自身的图像块嵌入将空间块聚为语义一致簇,通过掩码评估预测变化。CCI在忠实性基准上达到新高,例如在MS COCO检索任务中删除-AUC指标提升超一倍。结合GroundedSAM,CCI可自动分类预测为前景或背景驱动,具备关键诊断能力。现有基准如CounterAnimals仅以准确率衡量性能并默认错误全由背景相关导致,但分析表明视角变化、尺度偏移及细粒度混淆也贡献显著误差。为此,本文引入COVAR基准,系统性地分离前景与背景变量。借助CCI与COVAR,对18种CLIP变体进行综合评估,提供方法论进步与实证依据,推动更鲁棒的视觉语言模型发展。
原文摘要 · Abstract (English)
Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. We introduce Cluster-based Concept Importance (CCI), a novel interpretability method that uses CLIP's own patch embeddings to group spatial patches into semantically coherent clusters, mask them, and evaluate relative changes in model predictions. CCI sets a new state of the art on faithfulness benchmarks, surpassing prior methods by large margins; for example, it yields more than a twofold improvement on the deletion-AUC metric for MS COCO retrieval. We further propose that CCI, when combined with GroundedSAM, automatically categorizes predictions as foreground- or background-driven, providing a crucial diagnostic ability. Existing benchmarks such as CounterAnimals, however, rely solely on accuracy and implicitly attribute all performance degradation to background correlations. Our analysis shows this assumption to be incomplete, since many errors arise from viewpoint variation, scale shifts, and fine-grained object confusions. To disentangle these effects, we introduce COVAR, a benchmark that systematically varies object foregrounds and backgrounds. Leveraging CCI with COVAR, we present a comprehensive evaluation of eighteen CLIP variants, offering methodological advances and empirical evidence that chart a path toward more robust VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。