测试视觉语言模型在真实语义扰动下的表现,发现现有模型易受自然干扰影响。
Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks
- 构建对抗性数据集,评估模型在自然语义变化下的鲁棒性。
- 发现CLIP模型在自然语言诱导的对抗样本上性能下降超过40%。
- 揭示模型失败模式,为公平多模态识别提供新方向。
近年来,基于网络规模图像-文本对训练的视觉语言模型(VLMs)在多种视觉任务中实现了出色的零样本迁移能力。然而,全面且独立于标准基准的评估对于理解其鲁棒性、局限性和实际应用价值至关重要。本文提出一个系统性评估框架,针对多样化下游任务中的自然对抗场景对VLMs进行评测,这一方面此前被忽视。我们在典型字体攻击、ImageNet-A以及自然语言诱导的对抗样本等精心挑选的对抗数据集上,评估了多种VLMs(CLIP、robust CLIP、BLIP2和SigLIP2)在零样本图像分类、语义分割和视觉问答任务中的表现。分析显示,robust CLIP模型会放大自然对抗漏洞,而标准CLIP模型在自然语言诱导的对抗样本上性能显著下降。此外,我们提供了可解释性分析以识别模型的失效模式。期望这些发现能推动未来鲁棒且公平的多模态模式识别研究。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) trained on web-scale image-text pairs have enabled impressive zero-shot transfer across a diverse range of visual tasks. However, comprehensive and independent evaluation beyond standard benchmarks is essential to understand their robustness, limitations, and real-world applicability. This paper presents a systematic evaluation framework for VLMs under natural adversarial scenarios for diverse downstream tasks, which has been overlooked in previous evaluation works. We evaluate a wide range of VLMs (CLIP, robust CLIP, BLIP2, and SigLIP2) on curated adversarial datasets (typographic attacks, ImageNet-A, and natural language-induced adversarial examples). We measure the natural adversarial performance of selected VLMs for zero-shot image classification, semantic segmentation, and visual question answering. Our analysis reveals that robust CLIP models can amplify natural adversarial vulnerabilities, and CLIP models significantly reduce performance for natural language-induced adversarial examples. Additionally, we provide interpretable analyses to identify failure modes. We hope our findings inspire future research in robust and fair multimodal pattern recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。