arXiv:2502.01530cs.CVcs.CL2025-02被引 3

研究视觉语言模型在不同模态下的推理偏见差异

The in-context inductive biases of vision-language models differ across modalities

  • 通过三类实验对比模型在图文输入下的泛化倾向
  • 视觉输入时更倾向按形状分类,文本输入时形容词顺序影响判断
  • 不同模型表现不一,提示需关注输入模态对结果的影响

归纳偏置使学习者能在缺乏充分证据时做出推断。本文借鉴认知科学中的类别泛化范式,研究基础模型在上下文学习中的归纳偏置。现代基础模型可同时处理视觉与文本信息,其对不同模态的解读方式差异是新兴研究方向。我们通过三种实验范式,在三个视觉语言模型上考察刺激呈现模态及文本描述方式对泛化的影响。结果发现,模型普遍表现出对形状而非颜色的偏好;当示例以视觉形式呈现时,该形状偏置被强化;而当示例以文本形式呈现时,形容词顺序显著影响泛化结果。但这些效应在不同模型和实验范式间存在差异。研究揭示了视觉语言模型在上下文中对不同类型输入的表征机制,对实际应用具有指导意义。

原文摘要 · Abstract (English)

Inductive biases are what allow learners to make guesses in the absence of conclusive evidence. These biases have often been studied in cognitive science using concepts or categories -- e.g. by testing how humans generalize a new category from a few examples that leave the category boundary ambiguous. We use these approaches to study generalization in foundation models during in-context learning. Modern foundation models can condition on both vision and text, and differences in how they interpret and learn from these different modalities is an emerging area of study. Here, we study how their generalizations vary by the modality in which stimuli are presented, and the way the stimuli are described in text. We study these biases with three different experimental paradigms, across three different vision-language models. We find that the models generally show some bias towards generalizing according to shape over color. This shape bias tends to be amplified when the examples are presented visually. By contrast, when examples are presented in text, the ordering of adjectives affects generalization. However, the extent of these effects vary across models and paradigms. These results help to reveal how vision-language models represent different types of inputs in context, and may have practical implications for the use of vision-language models.

视觉语言模型归纳偏置多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。