arXiv:2506.00479cs.CLcs.CV2025-06ACL被引 6

构建首个统一评估无训练加速技术的基准,助力高效视觉语言模型落地

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models

  • 按令牌与参数压缩分类,系统评测主流无训练加速方法
  • 提出多维度评估框架,涵盖性能、泛化性与忠诚度等指标
  • 开源代码与方案,支持后续研究快速复现与优化

大型视觉语言模型(LVLMs)虽取得显著成果,但其高昂的计算开销制约了实际部署。尽管提升效率的研究不断涌现,现有方法在不同主干网络、评测基准和指标上的综合评估仍显不足。本文系统评估主流的LVLM加速技术,按令牌压缩与参数压缩两类进行划分。提出EffiVLM-Bench——一个统一评估框架,不仅衡量绝对性能,还关注泛化能力与模型忠实度,并探索帕累托最优权衡。通过大规模实验与深入分析,揭示了高效加速策略的潜在方向。相关代码与实现配方已开源,以推动未来研究。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable success, yet their significant computational demands hinder practical deployment. While efforts to improve LVLM efficiency are growing, existing methods lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. In this work, we systematically evaluate mainstream acceleration techniques for LVLMs, categorized into token and parameter compression. We introduce EffiVLM-Bench, a unified framework for assessing not only absolute performance but also generalization and loyalty, while exploring Pareto-optimal trade-offs. Our extensive experiments and in-depth analyses offer insights into optimal strategies for accelerating LVLMs. We open-source code and recipes for EffiVLM-Bench to foster future research.

视觉语言模型模型加速基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。