语言能动态重构视觉模型中的图像表示,提升任务相关性。
Linguistic Context Recodes Visual Representations in Vision-Language Models

- 用语言提示触发抽象的视觉目标表征,跨对象跨场景泛化。
- 后期层选择性增强目标相关属性,显著影响模型决策分布。
- 揭示视觉表示非静态,而是被语言动态调制,适合理解多模态机制的研究者。
目标导向的视觉处理是人类视觉智能的核心特征,生成支持分类或搜索等下游任务的表征。尽管视觉语言模型(VLMs)常面临类似任务,但其在语言提示下对视觉表征的动态重构能力仍不清晰。以往研究多将视觉表征视为静态信息库,由语言表征进行操控。本文首次提供两个具体证据:第一,识别出一种抽象参考表征,用于表示自然语言提示下的目标相关物体;通过提取对比性调控向量,证明其在模型预测中具有因果作用,且该表征具备跨物体、跨任务、跨合成与真实图像的泛化能力。第二,发现语言诱导的属性调制现象:后期层会针对性增强目标相关属性。该现象在多种提示下均成立。最后,通过因果干预验证属性调制主导了模型响应分布。结果支持更动态的跨模态处理观——视觉令牌并非静态信息库,而是被语言调制以响应语言提出的查询。
原文摘要 · Abstract (English)
Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their ability to recode visual representations when presented with goal-directed language remains poorly characterized. Indeed, prior work largely treats visual representations in VLMs as static repositories of visual information that are manipulated by language representations. In the present work, we provide evidence for two concrete instances of language-induced recoding of visual representations. First, we identify an abstract reference representation that denotes which objects are goal-relevant under a natural language prompt. We extract contrastive steering vectors corresponding to this reference representation and demonstrate that they are causally implicated in model predictions. These reference representations are abstract in that they generalize to different objects, different task contexts, and even from synthetic to naturalistic images. Second, we demonstrate language-induced attribute modulation: later layers selectively amplify goal-relevant attributes in visual representations of objects. We demonstrate this phenomenon across a range of different prompts. Finally, we provide a causal intervention that demonstrates that attribute modulation mediates a VLM's response distribution. Together, our results support a more dynamic account of cross-modality processing in VLMs -- rather than vision tokens serving as static repositories of information, they are modulated to support queries articulated in language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。