arXiv:2603.14989cs.CV2026-03被引 3

首个视觉语言模型推测解码基准,揭示多模态场景下加速新规律。

MMSpec: Benchmarking Speculative Decoding for Vision-Language Models

  • 构建统一框架评估10种推测解码算法在多模态任务中的表现。
  • 发现文本模型方法在多模态中性能下降,大批次下视觉感知更关键。
  • 提出ViSkip动态适配视觉令牌,实现最优吞吐与延迟平衡。

视觉语言模型(VLMs)在多模态任务中表现强劲,但因模型规模大、上下文长导致推理延迟高。推测解码作为新兴加速技术,其在VLMs中的表现仍不清晰。本文提出MMSpec,首个针对视觉语言模型的推测解码基准,包含600个跨六类任务的多模态样本,并集成十种代表性推测解码算法于统一评估框架。研究发现:(1) 专为纯文本大模型设计的方法在多模态场景中性能退化;(2) 大批量时视觉感知重要性显著提升;(3) 吞吐速度提升不能可靠反映延迟改善。基于此,提出ViSkip——一种即插即用的推测解码方法,动态适配视觉令牌,达到当前最优性能。

原文摘要 · Abstract (English)

Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculative decoding has recently emerged as an effective acceleration technique, yet its behavior in VLMs remains insufficiently understood. We introduce MMSpec, the first benchmark for evaluating speculative decoding in vision-language models. MMSpec contains 600 multimodal samples across six task categories and integrates ten representative speculative decoding algorithms under a unified evaluation framework. Our study reveals three key findings: (1) methods designed for text-only LLMs degrade in multimodal scenarios, (2) vision awareness becomes increasingly important at larger batch sizes, and (3) throughput speedup alone does not reliably reflect latency performance. Motivated by these findings, we propose ViSkip, a plug-and-play speculative decoding method that dynamically adapts speculation to vision tokens and achieves state-of-the-art performance.

视觉语言模型推测解码多模态推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。