arXiv:2501.05179cs.CV2025-01AAAI被引 31

通过全局缩略图指导局部图像压缩,显著提升高分辨率视觉语言模型推理效率

Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models

  • 用全局缩略图作为指挥官,动态引导局部图像的冗余去除
  • 压缩90%视觉令牌,性能保持90%以上,计算量降至9.1%
  • 适用于需要高效推理的高分辨率多视图视觉语言模型

大型视觉语言模型(LVLMs)在视觉理解方面表现优异,但因处理长多模态上下文时存在二次复杂度而面临效率挑战。现有令牌压缩方法针对单视图模型设计,未考虑高分辨率LVLM中动态裁剪带来的多视图特性。本文分析发现,全局缩略图可为局部裁剪提供整体上下文,用于判断信息量。基于此,提出新型即插即用压缩框架GlobalCom²,利用缩略图作为‘指挥官’,自适应保留关键细节并消除冗余。大量实验表明,GlobalCom²在压缩90%视觉令牌的同时,性能保持超过90%,显著降低计算量(至9.1%)和峰值内存(至60%)。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) excel at visual understanding, but face efficiency challenges due to quadratic complexity in processing long multi-modal contexts. While token compression can reduce computational costs, existing approaches are designed for single-view LVLMs and fail to consider the unique multi-view characteristics of high-resolution LVLMs with dynamic cropping. Existing methods treat all tokens uniformly, but our analysis reveals that global thumbnails can naturally guide the compression of local crops by providing holistic context for informativeness evaluation. In this paper, we first analyze dynamic cropping strategy, revealing both the complementary nature between thumbnails and crops, and the distinctive characteristics across different crops. Based on our observations, we propose ``Global Compression Commander'' (\textit{i.e.}, \textbf{GlobalCom$^2$}), a novel plug-and-play token compression framework for HR-LVLMs. GlobalCom$^2$ leverages thumbnail as the ``commander'' to guide the compression of local crops, adaptively preserving informative details while eliminating redundancy. Extensive experiments show that GlobalCom$^2$ maintains over \textbf{90\%} performance while compressing \textbf{90\%} visual tokens, reducing FLOPs and peak memory to \textbf{9.1\%} and \textbf{60\%}.

视觉语言模型令牌压缩推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。