解析视觉语言模型如何理解图像信息,揭示其处理机制。
Towards Interpreting Visual Information Processing in Vision-Language Models
- 通过消融实验分析视觉令牌在模型中的作用
- 移除特定物体令牌后识别准确率下降超70%
- 发现模型在最后层位置提取物体信息用于预测
视觉语言模型(VLMs)是处理和理解图文信息的强大工具。本文研究了代表性VLM LLaVA中语言模型组件对视觉令牌的处理过程,重点关注物体信息定位、视觉令牌表征在各层的演化以及视觉信息融合机制。通过消融实验发现,当移除与物体相关的视觉令牌时,物体识别准确率下降超过70%。研究观察到,随着网络层数加深,视觉令牌表征在词汇空间中逐渐变得可解释,显示出与图像内容对应的文本令牌之间的对齐趋势。此外,模型在最后一层的最后一个令牌位置提取经过优化的视觉信息以进行预测,这一机制与纯文本语言模型在事实关联任务中的行为一致。这些发现为理解VLM如何整合视觉信息提供了关键洞见,有助于弥合语言与视觉模型认知理解之间的差距,推动更可解释、可控的多模态系统发展。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses on analyzing the localization of object information, the evolution of visual token representations across layers, and the mechanism of integrating visual information for predictions. Through ablation studies, we demonstrated that object identification accuracy drops by over 70\% when object-specific tokens are removed. We observed that visual token representations become increasingly interpretable in the vocabulary space across layers, suggesting an alignment with textual tokens corresponding to image content. Finally, we found that the model extracts object information from these refined representations at the last token position for prediction, mirroring the process in text-only language models for factual association tasks. These findings provide crucial insights into how VLMs process and integrate visual information, bridging the gap between our understanding of language and vision models, and paving the way for more interpretable and controllable multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。