通过分析视觉标记的多层演化轨迹,实现高效低损的视觉令牌压缩。
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models

- 基于多层演化偏差评估视觉标记重要性,无需训练和注意力机制。
- 在保留11.1%视觉令牌时,保持94.4%的平均性能表现。
- 适合追求推理效率与精度平衡的大型视觉语言模型应用。
大型视觉语言模型(LVLM)在图像与视频理解任务中表现优异,但其推理效率受限于视觉编码器产生的大量视觉标记。现有压缩方法通常基于特定层的注意力分数或表示特性估算标记重要性,忽视了视觉标记在编码器各层间的演化过程。此类分层标准可能提供不完整的评估,限制压缩后的性能保持。我们分析了视觉标记在各层的演化方向,发现标记会形成多个群体演化路径。进一步观察表明,信息丰富的标记往往持续偏离这些群体演化方向。基于此,我们提出EvoCut,一种无需训练、无需注意力机制的视觉标记压缩方法,通过多层演化偏差评估标记重要性。实验结果表明,EvoCut在仅保留LLaVA-1.5-7B模型11.1%的视觉标记时,仍可保持94.4%的平均性能,充分验证了其在效率与准确性间的优越平衡能力。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) achieve strong performance on image and video understanding tasks, but their inference efficiency is constrained by the large number of visual tokens produced by vision encoders. Most existing visual token compression methods estimate token importance from attention scores or representation properties at specific layers, overlooking how visual tokens evolve across the vision encoder. Such layer-specific criteria may provide incomplete importance estimates and limit performance preservation after compression. To address this issue, we analyze layer-wise visual token evolution directions and observe that tokens form multiple group evolution directions across vision-encoder layers. Our analysis further shows that informative tokens tend to exhibit persistent deviations from common group evolution directions. Based on this observation, we propose EvoCut, a training-free and attention-free visual token compression method that estimates token importance from multi-layer evolution deviation. Experimental results show that EvoCut can retain only 11.1\% of the visual tokens on LLaVA-1.5-7B while preserving 94.4\% of the average performance, demonstrating its effectiveness in balancing efficiency and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。