arXiv:2504.13690cs.CV2025-04被引 18

首次系统分析视觉语言模型对19类常见图像损坏的鲁棒性差异。

Analysing the Robustness of Vision-Language-Models to Common Corruptions

  • 构建TextVQA-C与GQA-C新基准,评估文本识别与物体推理受干扰影响
  • 发现模型在模糊、雪灾下文本识别最脆弱,霜冻、脉冲噪声影响物体推理
  • 揭示变换器对低频特征偏好导致不同任务鲁棒性差异,指导抗干扰设计

视觉语言模型(VLMs)在理解视觉与文本内容方面展现出卓越能力,但其对常见图像损坏的鲁棒性仍缺乏系统研究。本文首次对来自ImageNet-C基准的19种损坏类型(涵盖噪声、模糊、天气和数字失真四类)进行全面分析。我们引入两个新基准TextVQA-C和GQA-C,分别用于系统评估损坏对场景文本理解与基于物体推理的影响。分析显示,基于Transformer的VLM在不同任务中表现出显著不同的脆弱模式:文本识别在模糊与雪灾损坏下退化最严重,而物体推理对霜冻与脉冲噪声更敏感。我们通过频率域特性关联这些现象,揭示了变换器对低频信息的固有偏好是造成差异化鲁棒性的关键原因。研究结果为开发更具抗干扰能力的视觉语言模型提供了重要洞见。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated impressive capabilities in understanding and reasoning about visual and textual content. However, their robustness to common image corruptions remains under-explored. In this work, we present the first comprehensive analysis of VLM robustness across 19 corruption types from the ImageNet-C benchmark, spanning four categories: noise, blur, weather, and digital distortions. We introduce two new benchmarks, TextVQA-C and GQA-C, to systematically evaluate how corruptions affect scene text understanding and object-based reasoning, respectively. Our analysis reveals that transformer-based VLMs exhibit distinct vulnerability patterns across tasks: text recognition deteriorates most severely under blur and snow corruptions, while object reasoning shows higher sensitivity to corruptions such as frost and impulse noise. We connect these observations to the frequency-domain characteristics of different corruptions, revealing how transformers' inherent bias toward low-frequency processing explains their differential robustness patterns. Our findings provide valuable insights for developing more corruption-robust vision-language models for real-world applications.

视觉语言模型鲁棒性分析图像损坏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。