全面评估CLIP模型在多维度下的鲁棒性,揭示其隐藏弱点与优化方向。
Toward a Holistic Evaluation of Robustness in CLIP Models
- 从视觉因子、模态对齐等六个维度系统测试CLIP鲁棒性
- 发现视觉编码器架构影响3D退化场景下的表现,且存在形状偏好
- 提示微调可缓解偏差,大模型融合能提升难分类任务性能
对比语言-图像预训练(CLIP)模型在零样本分类中展现出巨大潜力,尤其面对分布偏移时。本文旨在提供更全面的评估,引入多个新视角:首先,考察其对特定视觉因素变化的鲁棒性;其次,评估置信度不确定性与分布外检测这两个关键安全目标,超越单纯分类准确率;第三,衡量CLIP在图像与文本模态间对齐的精细程度;第四,扩展至3D感知能力,突破传统2D理解局限;第五,探究现代大模型(如LLaVA)中视觉-语言编码器交互对分类鲁棒性的影响。每个方面均分析六种因素:模型架构、训练分布、训练集规模、微调、对比损失和测试时提示。研究揭示若干此前未知的洞察:视觉编码器架构显著影响对3D退化场景的鲁棒性;CLIP模型预测常偏向形状特征;该偏差在ImageNet微调后减弱;利用CLIP视觉编码器的大模型在挑战类别上优于纯CLIP。结果为提升模型鲁棒性与可靠性提供重要指导。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-training (CLIP) models have shown significant potential, particularly in zero-shot classification across diverse distribution shifts. Building on existing evaluations of overall classification robustness, this work aims to provide a more comprehensive assessment of CLIP by introducing several new perspectives. First, we investigate their robustness to variations in specific visual factors. Second, we assess two critical safety objectives--confidence uncertainty and out-of-distribution detection--beyond mere classification accuracy. Third, we evaluate the finesse with which CLIP models bridge the image and text modalities. Fourth, we extend our examination to 3D awareness in CLIP models, moving beyond traditional 2D image understanding. Finally, we explore the interaction between vision and language encoders within modern large multimodal models (LMMs) that utilize CLIP as the visual backbone, focusing on how this interaction impacts classification robustness. In each aspect, we consider the impact of six factors on CLIP models: model architecture, training distribution, training set size, fine-tuning, contrastive loss, and test-time prompts. Our study uncovers several previously unknown insights into CLIP. For instance, the architecture of the visual encoder in CLIP plays a significant role in their robustness against 3D corruption. CLIP models tend to exhibit a bias towards shape when making predictions. Moreover, this bias tends to diminish after fine-tuning on ImageNet. Vision-language models like LLaVA, leveraging the CLIP vision encoder, could exhibit benefits in classification performance for challenging categories over CLIP alone. Our findings are poised to offer valuable guidance for enhancing the robustness and reliability of CLIP models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。