arXiv:2507.20198cs.CV2025-07综述被引 19

系统梳理多模态大模型的令牌压缩技术,助力高效处理长序列输入。

A Survey of Token Compression for Efficient Multimodal Large Language Models

  • 按视觉、视频、音频三类模态划分压缩方法,针对不同冗余特性设计策略
  • 提出四类核心机制:变换、相似性、注意力和查询驱动,覆盖主流压缩思路
  • 适合关注多模态模型效率优化的研究者与工程师快速入门

多模态大语言模型(MLLMs)在处理高分辨率图像、长视频序列和长音频输入等复杂上下文方面取得显著进展。然而,自注意力机制的二次方计算复杂度导致大量输入令牌带来巨大计算负担。为缓解这一瓶颈,令牌压缩成为关键且有前景的方法,在训练与推理中有效减少令牌数量。本文首次系统综述多模态长上下文令牌压缩领域的发展。鉴于有效压缩策略与各模态特性和冗余密切相关,我们按主要数据聚焦进行分类:(1) 图像中心压缩,应对视觉数据的空间冗余;(2) 视频中心压缩,解决动态序列中的时空冗余;(3) 音频中心压缩,处理声学信号的时间与频谱冗余。此外,从底层机制进一步分解为基于变换、相似性、注意力和查询的方法。本综述旨在整合当前进展,识别关键挑战,并激发该快速演进领域的未来研究方向。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have made remarkable strides, largely driven by their ability to process increasingly long and complex contexts, such as high-resolution images, extended video sequences, and lengthy audio input. While this ability significantly enhances MLLM capabilities, it introduces substantial computational challenges, primarily due to the quadratic complexity of self-attention mechanisms with numerous input tokens. To mitigate these bottlenecks, token compression has emerged as an auspicious and critical approach, efficiently reducing the number of tokens during both training and inference. In this paper, we present the first systematic survey and synthesis of the burgeoning field of multimodal long context token compression. Recognizing that effective compression strategies are deeply tied to the unique characteristics and redundancies of each modality, we categorize existing approaches by their primary data focus, enabling researchers to quickly access and learn methods tailored to their specific area of interest: (1) image-centric compression, which addresses spatial redundancy in visual data; (2) video-centric compression, which tackles spatio-temporal redundancy in dynamic sequences; and (3) audio-centric compression, which handles temporal and spectral redundancy in acoustic signals. Beyond this modality-driven categorization, we further dissect methods based on their underlying mechanisms, including transformation-based, similarity-based, attention-based, and query-based approaches. By providing a comprehensive and structured overview, this survey aims to consolidate current progress, identify key challenges, and inspire future research directions in this rapidly evolving domain.

多模态令牌压缩大模型效率综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。