arXiv:2608.20743cs.AI2026-08综述

提出统一框架,评估多模态模型在扩散并行生成中的可行性

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

论文配图:Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis
图 1 · 摘自论文原文
  • 构建分离并行度与树结构等设计的统一分类体系
  • 在OCR、VQA等多模态基准上实测不同并行策略效果
  • 揭示当前多模态并行生成的关键瓶颈与未来方向

推测解码通过轻量级草稿模型并行预估未来标记,由目标模型验证以加速自回归生成。其无损特性推动了草稿模型向并行生成演进。最新范式为块并行生成草稿,包括如DFlash和DSpark的扩散方法,在日常对话任务中实现最高3.6倍加速。尽管该趋势在纯文本大模型中已有深入研究,但其在多模态模型中的适用性仍不明确。现有工作聚焦输入压缩、适配器对齐、候选覆盖或模态特异性验证;然而,块并行生成草稿尚未被充分探索。为此,本文结合模态中心型综述与跨架构实证研究,探讨:多模态推测解码是否已准备好用于基于扩散的并行草稿?系统分析涵盖视觉-语言、视频-语言、音频及视觉-语言-动作(VLA)架构,从草稿并行性与跨模态信息交互双重视角出发。提出统一分类法,将草稿端并行性与树结构构造、验证策略等独立设计分离。进一步在标准化多模态基准(包括OCR、VQA、视觉推理与图像描述)上,针对不同并行程度进行综合比较。最后总结现有方法局限,讨论开放挑战,并展望该快速发展的领域未来路径。

原文摘要 · Abstract (English)

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.

多模态推测解码扩散模型并行生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。