arXiv:2510.04819cs.CVcs.CL2025-10中稿 · COLM被引 13

揭秘大模型如何处理视觉信息,发现其感知能力受限于关键值令牌的控制缺陷。

Visual Representations inside the Language Model

  • 分析多模态模型中视觉键值令牌的信息流动机制。
  • 发现语言模型在零样本任务中可完成分割、对应等感知任务,但信息量低于原始视觉编码器。
  • 提出通过文本前缀增强视觉表示,为改进模型感知提供新思路。

尽管已有大量关于ViT编码器和Transformer激活的可解释性研究,我们仍不清楚多模态语言模型(MLMs)为何在感知密集型任务上表现不佳。本文从一个被忽视的角度出发,考察主流MLMs(LLaVA-OneVision、Qwen2.5-VL、Llama-3-LLaVA-NeXT)如何处理其视觉键值令牌。首先研究视觉信息在语言模型中的传递路径,发现图像值令牌已包含足够信息,可在零样本条件下完成分割、语义对应、时间对应和指代表达检测等任务。虽然语言模型会对输入视觉编码的投影进行信息增强——该增强程度与整体感知能力相关——但在多项任务中,其视觉信息量仍低于未经过多模态微调的视觉编码器(SigLIP)。此外,发现语言模型后层中与输入无关的图像键令牌含有干扰性伪影,降低了整体感知能力。接着探讨对语言模型中视觉信息的调控,表明在图像输入前添加文本前缀可提升感知能力。最后揭示,若语言模型能更好控制视觉信息,其感知性能将显著提升:例如在BLINK基准的33.3%艺术风格问题中,语言模型内存在的感知信息并未传递至输出!本研究揭示了键值令牌在多模态系统中的作用,为深入理解多模态语言模型的机制可解释性提供了新方向,并提示训练视觉编码器与语言模型组件的新路径。

原文摘要 · Abstract (English)

Despite interpretability work analyzing VIT encoders and transformer activations, we don't yet understand why Multimodal Language Models (MLMs) struggle on perception-heavy tasks. We offer an under-studied perspective by examining how popular MLMs (LLaVA-OneVision, Qwen2.5-VL, and Llama-3-LLaVA-NeXT) process their visual key-value tokens. We first study the flow of visual information through the language model, finding that image value tokens encode sufficient information to perform several perception-heavy tasks zero-shot: segmentation, semantic correspondence, temporal correspondence, and referring expression detection. We find that while the language model does augment the visual information received from the projection of input visual encodings-which we reveal correlates with overall MLM perception capability-it contains less visual information on several tasks than the equivalent visual encoder (SigLIP) that has not undergone MLM finetuning. Further, we find that the visual information corresponding to input-agnostic image key tokens in later layers of language models contains artifacts which reduce perception capability of the overall MLM. Next, we discuss controlling visual information in the language model, showing that adding a text prefix to the image input improves perception capabilities of visual representations. Finally, we reveal that if language models were able to better control their visual information, their perception would significantly improve; e.g., in 33.3% of Art Style questions in the BLINK benchmark, perception information present in the language model is not surfaced to the output! Our findings reveal insights into the role of key-value tokens in multimodal systems, paving the way for deeper mechanistic interpretability of MLMs and suggesting new directions for training their visual encoder and language model components.

多模态可解释性视觉编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。