arXiv:2504.17892cs.CVcs.AI2025-04被引 6

提出一种自适应视觉标记压缩方法,提升多模态模型效率

Token Sequence Compression for Efficient Multimodal Computing

  • 通过聚类级标记聚合实现高效视觉信息压缩
  • 新方法在标记选择与合并上优于现有主流技术
  • 适合关注多模态模型加速与资源优化的研究者

大型多模态模型(LMMs)的迅猛发展推动了跨模态推理的进步,但带来了巨大的计算开销。本文聚焦于视觉语言模型,指出当前视觉编码器存在冗余与低效问题,提出一种自适应的多模态数据压缩方法。通过基准测试与定性分析,系统评估了多种视觉标记选择与合并策略。结果表明,简单的聚类级标记聚合在性能上超越了现有的最先进方法,包括在视觉编码器层级的合并及基于注意力的方法。研究揭示了当前视觉编码器中的冗余现象,并通过跨模态注意力可视化解释了若干令人困惑的视觉标记选择趋势。本工作是迈向更高效高维数据编码与处理的首次尝试,为构建可扩展、可持续的多模态系统铺平道路。

原文摘要 · Abstract (English)

The exponential growth of Large Multimodal Models (LMMs) has driven advancements in cross-modal reasoning but at significant computational costs. In this work, we focus on visual language models. We highlight the redundancy and inefficiency in current vision encoders, and seek to construct an adaptive compression method for multimodal data. In this work, we characterize a panoply of visual token selection and merging approaches through both benchmarking and qualitative analysis. In particular, we demonstrate that simple cluster-level token aggregation outperforms prior state-of-the-art works in token selection and merging, including merging at the vision encoder level and attention-based approaches. We underline the redundancy in current vision encoders, and shed light on several puzzling trends regarding principles of visual token selection through cross-modal attention visualizations. This work is a first effort towards more effective encoding and processing of high-dimensional data, and paves the way for more scalable and sustainable multimodal systems.

多模态视觉编码压缩高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。