arXiv:2511.13494cs.CVcs.AI2025-11被引 1

测试视觉语言模型对语义扰动的鲁棒性,发现大模型更稳定。

Language-Guided Invariance Probing of Vision-Language Models

  • 设计新评测基准,检测模型对同义改写和语义翻转的响应差异
  • 9个模型中,EVA02-CLIP和大版OpenCLIP表现最优,误差最低
  • 现有指标无法发现缺陷,该方法可诊断语言鲁棒性短板

近期视觉语言模型(如CLIP、OpenCLIP、EVA02-CLIP、SigLIP)在零样本任务中表现优异,但其对可控语言扰动的响应可靠性尚不明确。本文提出语言引导不变性探测(LGIP),评估模型在(i)语义保持的同义改写下的不变性,以及(ii)语义改变的语义翻转下的敏感性。基于4万张MS COCO图像及每图5个人工描述,自动生成同义改写与规则化翻转(涉及物体类别、颜色或数量变化),并用不变性误差、语义敏感度差距和正向率统计量总结模型行为。在九个VLM中,EVA02-CLIP和大型OpenCLIP变体位于有利的不变性-敏感性前沿,对同义改写方差小,且原始描述得分始终高于翻转版本。相比之下,SigLIP和SigLIP2的不变性误差显著更高,常偏好翻转描述,尤其在物体和颜色修改时。这些缺陷在标准检索指标下几乎不可见,表明LGIP能提供超越传统准确率的模型无关诊断,揭示视觉语言模型的语言鲁棒性问题。

原文摘要 · Abstract (English)

Recent vision-language models (VLMs) such as CLIP, OpenCLIP, EVA02-CLIP and SigLIP achieve strong zero-shot performance, but it is unclear how reliably they respond to controlled linguistic perturbations. We introduce Language-Guided Invariance Probing (LGIP), a benchmark that measures (i) invariance to meaning-preserving paraphrases and (ii) sensitivity to meaning-changing semantic flips in image-text matching. Using 40k MS COCO images with five human captions each, we automatically generate paraphrases and rule-based flips that alter object category, color or count, and summarize model behavior with an invariance error, a semantic sensitivity gap and a positive-rate statistic. Across nine VLMs, EVA02-CLIP and large OpenCLIP variants lie on a favorable invariance-sensitivity frontier, combining low paraphrase-induced variance with consistently higher scores for original captions than for their flipped counterparts. In contrast, SigLIP and SigLIP2 show much larger invariance error and often prefer flipped captions to the human descriptions, especially for object and color edits. These failures are largely invisible to standard retrieval metrics, indicating that LGIP provides a model-agnostic diagnostic for the linguistic robustness of VLMs beyond conventional accuracy scores.

视觉语言模型鲁棒性评测不变性探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。