arXiv:2602.15580cs.AI2026-02被引 5

揭示视觉如何在多模态模型中转化为语言,发现信息演化规律。

How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning

  • 用信息分解方法分析每层的视觉、语言和跨模态信息
  • 晚期语言信息占预测82%,视觉信息早期就衰减
  • 结果稳定且可解释,适合研究模型推理机制

当多模态Transformer回答视觉问题时,预测是源于视觉证据、语言推理,还是真正的跨模态融合?我们提出基于部分信息分解(PID)的逐层分析框架,将每层的预测信息分解为冗余、视觉独有、语言独有和协同成分。为使高维神经表征下的PID可行,引入 extit{PID Flow}:结合降维、归一化流高斯化与闭式高斯PID估计。在LLaVA-1.5-7B和LLaVA-1.6-7B上应用于六个GQA推理任务,发现一致的 extit{模态转换}模式:视觉独有信息早高峰后衰减,语言独有信息在深层激增,占最终预测约82\",跨模态协同始终低于2\"。该轨迹在模型变体间高度稳定(层间相关性>0.96),但强任务依赖,语义冗余决定信息指纹。通过定向图像→问题注意力切除实验验证因果性:破坏主转换路径导致视觉独有信息滞留增加、补偿性协同上升、总信息成本升高——这些效应在视觉依赖任务中最强,冗余任务中最弱。结果提供了多模态Transformer中视觉转语言的信息理论与因果解释,并为定位模态信息丢失的架构瓶颈提供量化指导。

原文摘要 · Abstract (English)

When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how does this structure evolve across layers? We address this question with a layer-wise framework based on Partial Information Decomposition (PID) that decomposes the predictive information at each Transformer layer into redundant, vision-unique, language-unique, and synergistic components. To make PID tractable for high-dimensional neural representations, we introduce \emph{PID Flow}, a pipeline combining dimensionality reduction, normalizing-flow Gaussianization, and closed-form Gaussian PID estimation. Applying this framework to LLaVA-1.5-7B and LLaVA-1.6-7B across six GQA reasoning tasks, we uncover a consistent \emph{modal transduction} pattern: visual-unique information peaks early and decays with depth, language-unique information surges in late layers to account for roughly 82\% of the final prediction, and cross-modal synergy remains below 2\%. This trajectory is highly stable across model variants (layer-wise correlations $>$0.96) yet strongly task-dependent, with semantic redundancy governing the detailed information fingerprint. To establish causality, we perform targeted Image$\rightarrow$Question attention knockouts and show that disrupting the primary transduction pathway induces predictable increases in trapped visual-unique information, compensatory synergy, and total information cost -- effects that are strongest in vision-dependent tasks and weakest in high-redundancy tasks. Together, these results provide an information-theoretic, causal account of how vision becomes language in multimodal Transformers, and offer quantitative guidance for identifying architectural bottlenecks where modality-specific information is lost.

多模态信息论模型解释视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。