arXiv:2603.11053cs.CLcs.IT2026-03

提出理论模型,提前预测推理加速的最优参数配置。

Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple

  • 建立预训练模型超参数与推理吞吐量的关系理论
  • 无需训练即可预测最佳推理配置,提升效率
  • 适合追求高效推理部署的研究者和工程师

推测解码是一种通过使用多个语言模型加速推理的技术。以往工作采用实验方式优化推理流水线的吞吐量,需进行大语言模型(LLM)训练,成本较高。本文研究推测解码,提出一种理论,可解析地将预训练LLM的关键超参数与下游基于推测解码的推理系统的吞吐量效率关联起来。该理论可在预训练前预测推理系统各组件的吞吐量最优超参数,从而实现更高效的推理部署。

原文摘要 · Abstract (English)

Speculative decoding is a technique that uses multiple language models to accelerate infer- ence. Previous works have used an experi- mental approach to optimize the throughput of the inference pipeline, which involves LLM training and can be costly. This study of spec- ulative decoding proposes a theory that ana- lytically connects the key hyperparameters of pre-trained LLMs to the throughput efficiency of a downstream SD-based inference system. The theory allows the prediction of throughput- optimal hyperparameters for the components of an inference system before their pre-training.

推理加速模型优化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。