arXiv:2511.14582cs.CV2025-11被引 24

用音频引导动态压缩音视频令牌,加速多模态大模型推理。

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

  • 通过音频显著性识别和保留分数,动态指导视频令牌裁剪。
  • 实现3.42倍推理速度提升,内存减少1.4倍,性能不降。
  • 无需训练,适合追求高效推理的多模态模型部署者。

多模态大语言模型(OmniLLMs)近年来受到广泛关注,致力于统一音频-视频理解。然而,处理更长的音视频联合令牌序列带来的高计算成本已成为关键瓶颈。现有令牌压缩方法未解决联合压缩多模态令牌的需求。为此,我们提出OmniZip,一种无需训练的、音频引导的音视频令牌压缩框架,优化多模态令牌表示并加速模型推理。具体而言,OmniZip首先识别显著性音频令牌,计算每个时间组的音频保留分数以捕捉信息密度,从而动态引导视频令牌剪枝,并通过跨模态相似性增强音频锚点线索。对于每个时间窗口,采用交错的时空方案压缩视频令牌。大量实验表明,OmniZip相比其他顶尖方法实现3.42倍推理加速与1.4倍内存降低,且无需训练即可保持OmniLLMs性能。

原文摘要 · Abstract (English)

Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However, the high computational cost of processing longer joint audio-video token sequences has become a key bottleneck. Existing token compression methods have not addressed the emerging need to jointly compress multimodal tokens. To bridge this gap, we present OmniZip, a training-free, audio-guided audio-visual token-compression framework that optimizes multimodal token representation and accelerates model inference. Specifically, OmniZip first identifies salient audio tokens, then computes an audio retention score for each time group to capture information density, thereby dynamically guiding video token pruning and preserving cues from audio anchors enhanced by cross-modal similarity. For each time window, OmniZip compresses the video tokens using an interleaved spatio-temporal scheme. Extensive results demonstrate the merits of OmniZip: it achieves a 3.42X inference speedup and a 1.4X memory reduction over other top-performing counterparts, while maintaining the performance of OmniLLMs without training.

多模态令牌压缩推理加速音频引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。