通过分析多模态大模型的隐藏表示,揭示不同融合方式对模型内部结构的影响。
MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models

- 设计新系统分析多模态大模型中跨层嵌入的线性、维度与各向异性
- OmniFusion图像嵌入维度更高且各向异性更稳定,而LLaVA-NeXT图像线性略降
- 发现模态融合方式直接影响模型内部结构,为未来设计提供依据
本文提出MLLM-Microscope,一个用于分析多模态大语言模型(MLLMs)隐藏表示的新系统。该系统评估了变压器层间多模态令牌嵌入的线性、内在维度和各向异性。基于ScienceQA数据集,我们评估了两种先进MLLMs:LLaVA-NeXT与OmniFusion。结果表明,两种模型中主路径与残差路径的双模态令牌均表现出高度线性行为。然而,LLaVA-NeXT的图像令牌线性略有下降,而OmniFusion保持稳定。此外,OmniFusion的图像令牌在各层中维度始终高于LLaVA-NeXT。OmniFusion的各向异性在整个层级中也保持较低水平。这些发现表明,MLLM内部运作高度依赖于模态融合方式。本系统所揭示的其他潜在洞见,有助于深化对MLLM内部机制的理解,推动未来模型设计与优化。
原文摘要 · Abstract (English)
This work presents MLLM-Microscope, a novel system designed for analyzing the hidden representations within Multimodal Large Language Models (MLLMs). Our system evaluates the linearity, intrinsic dimension, and anisotropy of multimodal token embeddings across transformer layers. Utilizing the ScienceQA dataset, we evaluate two state-of-the-art MLLMs, LLaVA-NeXT and OmniFusion. We find that both the main and residual streams for tokens of both modalities exhibit highly linear behaviors across transformer layers. However, LLaVA-NeXT's image tokens reveal a slight decline in linearity, whereas OmniFusion's remain consistent. Image token dimensions in OmniFusion remain consistently higher across layers compared to LLaVA-NeXT. Also, the OmniFusion's anisotropy is observed to stay consistently low throughout the layers. These findings suggest that the inner workings of MLLMs highly depend on the nature of modality fusion performed before passing the token sequence into LLM. This and other new potential insights obtainable from our system are surely capable of enhancing our understanding of the inner workings of MLLMs, informing future model design and optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。