arXiv:2503.13108cs.CVcs.AI2025-03CVPR被引 26

揭示多模态模型视觉信息处理路径,实现推理加速65%不降性能。

Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference

论文配图:Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
图 1 · 摘自论文原文
  • 发现视觉信息在浅层注入指令词,深层自交互优化视觉表征。
  • 提出分层感知剪枝法,特定层动态裁剪图像标记,降低65%计算量。
  • 适合追求高效推理的多模态模型开发者,尤其关注速度与精度平衡者。

多模态大语言模型(MLLMs)通过将预训练视觉编码器的视觉特征融合到大语言模型(LLMs)中,提升了视觉-语言任务的表现。然而,MLLMs如何处理和利用视觉信息仍不清晰。本文揭示了视觉信息流动的显著转变:(1) 在浅层,图像标记与指令标记间存在强交互,大部分视觉信息被注入指令标记以形成跨模态语义表示;(2) 在深层,图像标记主要相互交互,聚合剩余视觉信息以优化视觉模态内的语义表示。基于此洞察,我们提出分层模态感知剪枝(HiMAP),一种即插即用的推理加速方法,可在特定层动态剪枝图像标记,计算成本降低约65%,且不牺牲性能。研究结果为理解MLLMs中的视觉信息处理提供了新视角,并提供了当前最优的高效推理解决方案。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant flow of visual information is uncovered: (1) in shallow layers, strong interactions are observed between image tokens and instruction tokens, where most visual information is injected into instruction tokens to form cross-modal semantic representations; (2) in deeper layers, image tokens primarily interact with each other, aggregating the remaining visual information to optimize semantic representations within visual modality. Based on these insights, we propose Hierarchical Modality-Aware Pruning (HiMAP), a plug-and-play inference acceleration method that dynamically prunes image tokens at specific layers, reducing computational costs by approximately 65% without sacrificing performance. Our findings offer a new understanding of visual information processing in MLLMs and provide a state-of-the-art solution for efficient inference.

多模态推理加速视觉表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。