提出生成-精炼框架分类体系,加速自回归模型推理
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
- 按生成策略与精炼机制分类,系统梳理自回归生成方法
- 涵盖文本、图像、语音生成,支持多场景高效部署
- 适合关注大模型推理加速的研究者与工程开发者
自回归模型的序列依赖关系是大规模部署中的根本瓶颈,尤其在实时应用中。传统优化方法如剪枝和量化常以牺牲模型质量为代价,而近期生成-精炼框架的进展表明该权衡可显著缓解。本综述构建了生成-精炼框架的全面分类体系,分析了自回归序列任务中的各类方法。根据生成策略(从简单的n-gram预测到复杂的草稿模型)与精炼机制(包括单次验证与迭代方法)进行分类。通过系统分析算法创新与系统级实现,考察了不同计算环境下的部署策略,并探讨了文本、图像与语音生成等多领域应用。本研究为高效自回归解码的未来研究提供了理论与实践基础。
原文摘要 · Abstract (English)
Sequential dependencies present a fundamental bottleneck in deploying large-scale autoregressive models, particularly for real-time applications. While traditional optimization approaches like pruning and quantization often compromise model quality, recent advances in generation-refinement frameworks demonstrate that this trade-off can be significantly mitigated. This survey presents a comprehensive taxonomy of generation-refinement frameworks, analyzing methods across autoregressive sequence tasks. We categorize methods based on their generation strategies (from simple n-gram prediction to sophisticated draft models) and refinement mechanisms (including single-pass verification and iterative approaches). Through systematic analysis of both algorithmic innovations and system-level implementations, we examine deployment strategies across computing environments and explore applications spanning text, images, and speech generation. This systematic examination of both theoretical frameworks and practical implementations provides a foundation for future research in efficient autoregressive decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。