研究视觉语言模型在模态冲突时如何选择信息来源。
How Do Vision-Language Models Process Conflicting Information Across Modalities?
- 通过不一致图文配对测试模型的模态偏好。
- 发现不同模型倾向图像或文本,且内部注意力头可调控偏好。
- 提出可转移的路由头,提升跨数据集多模态表现。
随着AI系统日益需要处理多模态输入,如何将不同来源的信息整合为连贯的表征成为关键问题。本文聚焦视觉语言模型,在输入存在矛盾时(如一张狗的图片配以‘猫的照片’的文字描述),考察模型对特定模态信息的响应能力。实验发现,多数模型倾向于优先采纳图像信息,但不同模型的偏好存在差异。进一步分析显示,这种行为偏好反映在模型内部表示结构中,特定注意力头可重构表示以强化某一模态。此外,我们识别出一类模态无关的‘路由头’,其能根据指令偏好引导模型回答对应模态内容,且该机制可被操控或迁移,从而提升在不同数据集和模态上的性能。本工作为识别与控制模型在复杂多模态环境中的冲突信号处理机制提供了基础步骤。
原文摘要 · Abstract (English)
AI models are increasingly required to be multimodal, integrating disparate input streams into a coherent state representation on which subsequent behaviors and actions can be based. This paper seeks to understand how such models behave when input streams present conflicting information. Focusing specifically on vision-language models, we provide inconsistent inputs (e.g., an image of a dog paired with the caption "A photo of a cat") and ask the model to report the information present in one of the specific modalities (e.g., "What does the caption say / What is in the image?"). We find that models often favor one modality over the other, e.g., reporting the image regardless of what the caption says, but that different models differ in which modality they favor. We find evidence that the behaviorally preferred modality is evident in the internal representational structure of the model, and that specific attention heads can restructure the representations to favor one modality over the other. Moreover, we find modality-agnostic "router heads" which appear to promote answers about the modality requested in the instruction, and which can be manipulated or transferred in order to improve performance across datasets and modalities. Together, the work provides essential steps towards identifying and controlling if and how models detect and resolve conflicting signals within complex multimodal environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。