建立首个统一的视觉标记压缩评估基准,提升多模态模型推理效率。
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
- 构建统一评测框架,覆盖六维度十数据集
- 发现随机剪枝是强基线,压缩率决定性能下降
- OCR任务最敏感,无万能压缩方法
大型多模态模型(LMMs)因图像编码器引入大量视觉标记而存在严重推理效率问题。尽管近期剪枝与合并等标记压缩方法显示出减少冗余的潜力,其评估仍碎片化且不一致。本文提出UniPruneBench,一个统一可扩展的多模态大模型视觉标记剪枝基准。该基准涵盖六个能力维度和十个数据集,覆盖十种代表性压缩算法及三类LMMs(LLaVA-v1.5、Intern-VL3、Qwen2.5-VL)。除任务准确率外,还纳入运行时间、预填充延迟等系统级指标,提供全面评估。实验揭示:(1) 随机剪枝是出人意料的强基线;(2) 无单一方法在所有场景中持续领先;(3) 剪枝敏感性任务间差异显著,OCR最脆弱;(4) 压缩比例是决定性能退化的主导因素。我们相信UniPruneBench将为未来高效多模态建模研究提供可靠基础。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) often suffer from severe inference inefficiency due to the large number of visual tokens introduced by image encoders. While recent token compression methods, such as pruning and merging, have shown promise in reducing redundancy, their evaluation remains fragmented and inconsistent. In this work, we present UniPruneBench, a unified and extensible benchmark for visual token pruning in multimodal LLMs. UniPruneBench provides standardized protocols across six ability dimensions and ten datasets, covering ten representative compression algorithms and three families of LMMs (LLaVA-v1.5, Intern-VL3, and Qwen2.5-VL). Beyond task accuracy, it incorporates system-level metrics such as runtime and prefilling latency to provide a holistic view. Our experiments uncover several key findings: (1) random pruning is a surprisingly strong baseline, (2) no single method consistently outperforms others across scenarios, (3) pruning sensitivity varies significantly across tasks, with OCR being most vulnerable, and (4) pruning ratio is the dominant factor governing performance degradation. We believe UniPruneBench will serve as a reliable foundation for future research on efficient multimodal modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。