arXiv:2606.20077cs.CVcs.AI2026-06

对比视觉信息在大模型中的两种融合方式,揭示其内部演化差异。

The Hidden Evolution of Disguised Visual Context inside the VLM

论文配图:The Hidden Evolution of Disguised Visual Context inside the VLM
图 1 · 摘自论文原文
  • 通过统一训练条件,比较输入层与中间层注入的视觉融合策略。
  • 发现视觉信号在模型中被逐步重构,不同方式捕捉不同频段特征。
  • 性能差异源于表征质量而非注意力分配,适合多模态研究者参考。

视觉标记以原始、非语言信号形式进入大型语言模型(LLMs)。它们如何转化为有意义的表示并融入语言空间,完全取决于集成架构——或作为输入序列中的上下文提示,或直接注入到LLM的中间层。当前对这两种架构选择如何影响视觉信息及其内部转化过程的理解仍不充分。我们通过在单图、多图和视频基准上,在相同训练条件下,公平比较基于上下文和分层注入的VLM集成范式。结果揭示了一种隐藏的演化:视觉标记以伪装的视觉上下文形式进入LLM,即缺乏语言结构的原始表征,但会根据集成方式逐步重塑,各自捕获视觉信号的根本不同频率特性。我们证明,这种内部演化决定了VLM能有效利用的视觉特征、视觉表示与语言空间的对齐程度,以及各范式在不同任务上的表现。进一步表明,仅靠注意力分配不足,性能由每层视觉表征的质量决定。

原文摘要 · Abstract (English)

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visual tokens as in-context prompts within the input sequence or injecting them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect visual information and its internal transformation to integrate with the LLM remains underexplored. We provide a fair comparison by evaluating in-context and layer-wise injection VLM integration paradigms under identical training conditions across single image, multi-image, and video benchmarks. In doing so, we uncover a hidden evolution where visual tokens enter the LLM as disguised visual context, raw representations lacking linguistic structure, but are progressively reshaped depending on the integration paradigm, each capturing fundamentally different frequency characteristics of the visual signal. We show that this evolution inside the LLM determines what visual features the VLM can utilize effectively, how visual representations align with the language space, and ultimately how each paradigm performs across different tasks. We further demonstrate that attention allocation alone is insufficient, and that performance is driven by the quality of visual representations at each layer.

多模态视觉表征大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。