arXiv:2412.06646cs.CVcs.LG2024-12NeurIPS被引 4

发现视觉信息在联合生成模型中通过单一关键令牌传递,可精准控制图文内容。

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

  • 对比原生与非原生多模态模型,发现前者依赖单一视觉通路令牌
  • 删除该令牌导致图像理解性能显著下降,证明其核心作用
  • 可通过精准干预该令牌实现细粒度图文语义调控,适合可控生成研究者

近期多模态训练进展显著提升了统一模型中图像理解与生成的融合能力。本文研究视觉语言模型(VLMs)如何处理图像理解任务,重点分析视觉信息在模型中的处理与传递机制。我们比较了原生多模态VLMs(从头训练于多模态数据,可生成文本和图像)与非原生多模态VLMs(由预训练语言模型适配或仅支持文本生成),揭示两者在信息流上的关键差异。结果表明,在原生多模态模型中,图像与文本嵌入在残差流中更为分离;视觉信息传至文本的方式也不同:非原生模型呈现分布式通信模式,信息通过多个图像令牌传递;而原生联合生成模型则倾向于依赖单个图像后令牌作为视觉信息的‘窄门’。我们证明,删除该令牌会显著损害图像理解性能,而针对性的令牌级干预能可靠地引导图像语义与下游文本生成,实现细粒度控制。

原文摘要 · Abstract (English)

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-language models (VLMs) handle image-understanding tasks, focusing on how visual information is processed and transferred to the textual domain. We compare native multimodal VLMs, models trained from scratch on multimodal data to generate both text and images, and non-native multimodal VLMs, models adapted from pre-trained large language models or capable of generating only text, highlighting key differences in information flow. We find that in native multimodal VLMs, image and text embeddings are more separated within the residual stream. Moreover, VLMs differ in how visual information reaches text: non-native multimodal VLMs exhibit a distributed communication pattern, where information is exchanged through multiple image tokens, whereas models trained natively for joint image and text generation tend to rely on a single post-image token that acts as a narrow gate for visual information. We show that ablating this single token significantly deteriorates image-understanding performance, whereas targeted, token-level interventions reliably steer image semantics and downstream text with fine-grained control.

多模态模型视觉通路可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。