arXiv:2510.07143cs.CV2025-10ACL被引 12

现有视觉压缩评估存在偏差,新框架通过降采样过滤噪声提升公平性。

Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods

  • 用图像降采样作为筛选器,识别压缩敏感样本
  • 实验证明降采样比多数先进压缩方法更有效
  • 提出VTC-Bench框架,解决评估基准与任务不匹配问题

近期加速多模态大模型推理的研究主要聚焦于视觉标记压缩。当前方法通常通过压缩前后在现有多模态大模型基准上的准确率下降来评估效果。然而这些基准原本设计用于评估通用感知与推理能力,而非针对视觉标记压缩的特定挑战,导致任务与评估不匹配。本文揭示了一个反直觉但一致的现象:简单图像降采样在多个主流基准上优于许多先进的视觉标记压缩方法。通过涵盖八个常用基准和多种前沿压缩技术的全面实证研究,我们发现(i)现有基准包含大量与任务无关的噪声样本,影响压缩评估;(ii)降采样可作为有效数据筛选器,区分压缩敏感与非敏感样本。基于此,我们提出VTC-Bench评估框架,明确利用降采样作为判别器,对现有基准去噪,实现对视觉标记压缩方法更公平、更有意义的额外评估。

原文摘要 · Abstract (English)

Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly evaluated by measuring the accuracy drop on existing MLLM benchmarks before and after compression. However, these benchmarks are originally designed to assess general perception and reasoning abilities, rather than the specific challenges posed by visual token compression, leading to a fundamental task mismatch. In this work, we uncover a counterintuitive yet consistent phenomenon: simple image downsampling outperforms many advanced visual token compression methods across multiple widely used benchmarks. Through a comprehensive empirical study spanning eight popular benchmarks and multiple state-of-the-art compression techniques, we show that (i) current benchmarks contain substantial noise (task-irrelevant samples) for evaluating visual token compression, and (ii) downsampling can act as an effective data filter that distinguishes between simple and difficult samples with respect to compression sensitivity. Motivated by these findings, we propose VTC-Bench, an evaluation framework that explicitly leverages downsampling as a discriminator to denoise existing benchmarks, enabling a fairer and more meaningful additional assessment of visual token compression methods.

视觉压缩评估框架多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。