arXiv:2509.23050cs.LGcs.AI2025-09被引 13

通过嵌入链分析,揭示视觉如何在特定层影响大模型决策。

Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding

  • 用嵌入链追踪模型各层表征变化,定位视觉信息介入的关键层
  • 发现所有模型都有视觉整合点(VIP),且其后视觉影响显著增强
  • 提出TVI指标,可量化语言先验强度,适合模型调试与评估

大型多模态模型虽在任务中表现优异,但常依赖预训练中记忆的文本模式(语言先验),忽视视觉信息。现有分析依赖输入输出探查,无法揭示内部机制。本文首次通过嵌入链分析,研究多模态模型中表征的逐层演化。结果发现:几乎所有模型均存在一个视觉整合点(VIP),即视觉信息开始显著重塑隐藏表征并影响生成决策的关键层。基于此,我们提出总视觉整合度(TVI)指标,衡量越过VIP后的表征差异,以量化视觉查询对输出的影响强度。在10个主流模型与6个基准数据集共60种组合上验证,VIP普遍出现,且TVI能可靠预测语言先验强弱。该工作为诊断和理解多模态模型中的语言先验提供了系统性工具。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) achieve strong performance on multimodal tasks, yet they often default to their language prior (LP) -- memorized textual patterns from pre-training while under-utilizing visual evidence. Prior analyses of LP mostly rely on input-output probing, which fails to reveal the internal mechanisms governing when and how vision influences model behavior. To address this gap, we present the first systematic analysis of language prior through the lens of chain-of-embedding, which examines the layer-wise representation dynamics within LVLMs. Our analysis reveals a universal phenomenon: each model exhibits a Visual Integration Point (VIP), a critical layer at which visual information begins to meaningfully reshape hidden representations and influence decoding for multimodal reasoning. Building on this observation, we introduce the Total Visual Integration (TVI) estimator, which aggregates representational discrepancy beyond the VIP to quantify how strongly visual query influences response generation. Across 60 model-dataset combinations spanning 10 contemporary LVLMs and 6 benchmarks, we demonstrate that VIP consistently emerges, and that TVI reliably predicts the strength of language prior. This offers a principled toolkit for diagnosing and understanding language prior in LVLMs.

多模态模型语言先验嵌入链视觉整合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。