arXiv:2411.00841cs.LGcs.AI2024-11NeurIPS被引 33

从马尔可夫链视角解析大模型推测解码的理论极限与效率机制

A Theoretical Perspective for Speculative Decoding Algorithm

  • 用马尔可夫链建模推测解码过程,揭示其核心机制
  • 证明了输出质量与推理加速之间的理论权衡边界
  • 为小模型生成、大模型验证的协同机制提供数学依据

基于Transformer的自回归采样是大语言模型推理速度的主要瓶颈。推测解码是一种有效加速方法,通过小模型生成候选词序列,再由大模型进行验证。尽管其在实践中表现良好,但其理论理解仍显不足。本文通过马尔可夫链抽象化建模解码过程,从理论上分析了推测解码的输出质量与推理加速两大核心属性。研究覆盖了推测解码的理论极限、批处理算法以及质量与加速间的权衡关系。结果揭示了不同组件间通过总变差距离建立的根本联系,并说明这些因素如何共同影响解码效率。

原文摘要 · Abstract (English)

Transformer-based autoregressive sampling has been the major bottleneck for slowing down large language model inferences. One effective way to accelerate inference is \emph{Speculative Decoding}, which employs a small model to sample a sequence of draft tokens and a large model to validate. Given its empirical effectiveness, the theoretical understanding of Speculative Decoding is falling behind. This paper tackles this gap by conceptualizing the decoding problem via markov chain abstraction and studying the key properties, \emph{output quality and inference acceleration}, from a theoretical perspective. Our analysis covers the theoretical limits of speculative decoding, batch algorithms, and output quality-inference acceleration tradeoffs. Our results reveal the fundamental connections between different components of LLMs via total variation distances and show how they jointly affect the efficiency of decoding algorithms.

大模型推理推测解码理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。