arXiv:2509.12492cs.CV2025-09被引 1

测试视觉语言模型在噪声下的表现,发现模型大小与抗噪能力无简单关系。

Evaluating Robustness of Vision-Language Models Under Noisy Conditions

  • 在光照、模糊、压缩等噪声下评估多个主流视觉语言模型。
  • 大模型如LLaVA语义理解强,但不总比小模型抗噪能力强。
  • JPEG压缩和运动模糊会显著降低各类模型性能。

视觉语言模型(VLMs)在图像描述和视觉问答等多模态任务中取得卓越进展,但在噪声条件下的鲁棒性仍不明确。本研究提出一个全面评估框架,在受控扰动(包括光照变化、运动模糊和压缩伪影)下评估多个前沿VLMs的表现。采用基于词汇的指标(BLEU、METEOR、ROUGE、CIDEr)和基于句子嵌入的神经相似度度量来量化语义对齐。实验覆盖多个数据集,揭示关键发现:(1) 真实标注的描述丰富程度显著影响模型性能;(2) 更大的模型如LLaVA在语义理解上表现优异,但并非普遍优于小模型;(3) 某些噪声类型(如JPEG压缩和运动模糊)会大幅削弱各类模型性能。研究结果揭示了模型规模、数据集特性与噪声鲁棒性之间的复杂权衡,为未来鲁棒多模态学习提供了标准化基准。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have attained exceptional success across multimodal tasks such as image captioning and visual question answering. However, their robustness under noisy conditions remains unfamiliar. In this study, we present a comprehensive evaluation framework to evaluate the performance of several state-of-the-art VLMs under controlled perturbations, including lighting variation, motion blur, and compression artifacts. We used both lexical-based metrics (BLEU, METEOR, ROUGE, CIDEr) and neural-based similarity measures using sentence embeddings to quantify semantic alignment. Our experiments span diverse datasets, revealing key insights: (1) descriptiveness of ground-truth captions significantly influences model performance; (2) larger models like LLaVA excel in semantic understanding but do not universally outperform smaller models; and (3) certain noise types, such as JPEG compression and motion blur, dramatically degrade performance across models. Our findings highlight the nuanced trade-offs between model size, dataset characteristics, and noise resilience, offering a standardized benchmark for future robust multimodal learning.

视觉语言模型鲁棒性评估噪声干扰

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。