研究视觉模型在异常图像下的理解偏差,发现语义与图文一致性的冲突会削弱模型推理能力。
VISaGE: Understanding Visual Generics and Exceptions
- 构建新数据集VISaGE,对比典型与异常图像的模型表现
- 图文不一致时,模型概念理解能力显著下降
- 适合关注模型鲁棒性与推理机制的研究者
尽管视觉语言模型(VLMs)在训练中学习了泛化知识,但通常仅用于分析单个实例。当评估样本具有异常性时,模型中的两种先验产生冲突:一是基于微调数据的图文一致性实用先验;二是类别概念普遍成立的语义先验。为理解模型如何权衡这两种先验,我们引入新评估数据集VISaGE,包含典型与异常图像。在精心设计的实验中,我们发现当图文不一致时,模型的语义理解能力显著下降,且该影响强于语义先验对个体实例的影响。
原文摘要 · Abstract (English)
While Vision Language Models (VLMs) learn conceptual representations, in the form of generalized knowledge, during training, they are typically used to analyze individual instances. When evaluation instances are atypical, this paradigm results in tension between two priors in the model. The first is a pragmatic prior that the textual and visual input are both relevant, arising from VLM finetuning on congruent inputs; the second is a semantic prior that the conceptual representation is generally true for instances of the category. In order to understand how VLMs trade off these priors, we introduce a new evaluation dataset, VISaGE, consisting of both typical and exceptional images. In carefully balanced experiments, we show that conceptual understanding degrades when the assumption of congruency underlying the pragmatic prior is violated with incongruent images. This effect is stronger than the effect of the semantic prior when querying about individual instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。