用信息分解法解析视觉语言模型决策机制,揭示融合与依赖的真相
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
- 基于部分信息分解,量化模型决策中冗余、独立和协同的信息成分
- 发现两种任务模式和两类模型家族策略,层间处理有三阶段规律
- 适合研究多模态模型机制、评估模型真实融合能力的研究者
大型视觉语言模型(LVLMs)表现优异,但其内部决策过程不透明,难以判断成功是源于真正的多模态融合,还是依赖单模态先验。为此,我们提出一种基于部分信息分解(PID)的新框架,定量分析LVLM的“信息谱”——将决策相关信息分解为冗余、独特和协同三部分。通过适配可扩展估计器至现代LVLM输出,我们的模型无关方法在四个数据集上对26个LVLM进行了多维度分析:广度(跨模型与跨任务)、深度(逐层信息动态)和时间(训练过程中的动态变化)。结果揭示两个关键发现:(i) 两种任务范式(协同驱动型与知识驱动型);(ii) 两类稳定且相反的家族级策略(融合中心型与语言中心型)。还发现层间处理存在一致的三阶段模式,并识别出视觉指令微调是融合学习的关键阶段。这些成果提供了超越准确率的定量分析视角,为下一代LVLM的设计与分析提供洞见。代码与数据见https://github.com/RiiShin/pid-lvlm-analysis。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) achieve impressive performance, yet their internal decision-making processes remain opaque, making it difficult to determine if the success stems from true multimodal fusion or from reliance on unimodal priors. To address this attribution gap, we introduce a novel framework using partial information decomposition (PID) to quantitatively measure the "information spectrum" of LVLMs -- decomposing a model's decision-relevant information into redundant, unique, and synergistic components. By adapting a scalable estimator to modern LVLM outputs, our model-agnostic pipeline profiles 26 LVLMs on four datasets across three dimensions -- breadth (cross-model & cross-task), depth (layer-wise information dynamics), and time (learning dynamics across training). Our analysis reveals two key results: (i) two task regimes (synergy-driven vs. knowledge-driven) and (ii) two stable, contrasting family-level strategies (fusion-centric vs. language-centric). We also uncover a consistent three-phase pattern in layer-wise processing and identify visual instruction tuning as the key stage where fusion is learned. Together, these contributions provide a quantitative lens beyond accuracy-only evaluation and offer insights for analyzing and designing the next generation of LVLMs. Code and data are available at https://github.com/RiiShin/pid-lvlm-analysis .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。