探究大模型推理中推测解码的不公平加速现象
The Disparate Impacts of Speculative Decoding
- 分析推测解码在不同任务上的加速差异
- 发现低资源任务平均提速仅12%提升
- 提出缓解策略,适合优化推理公平性的研究者
推测解码通过使用更小、更便宜的「起草模型」来概率性支持推理,已成为系统降低大语言模型解码时间的标准技术。本文从任务间加速率差异的角度分析推测解码的影响,发现其加速效果并非均匀分布,对欠拟合及常被忽视的任务而言,加速效果显著减弱。论文推导出量化此‘不公平性’的分析框架,并揭示导致加速差异的驱动因素。基于这些洞察,提出一种缓解策略,在多个模型对上验证,平均使公平性指标提升12%。
原文摘要 · Abstract (English)
The practice of speculative decoding, whereby inference is probabilistically supported by a smaller, cheaper, ``drafter'' model, has become a standard technique for systematically reducing the decoding time of large language models. This paper conducts an analysis of speculative decoding through the lens of its potential disparate speed-up rates across tasks. Crucially, the paper shows that speed-up gained from speculative decoding is not uniformly distributed across tasks, consistently diminishing for under-fit, and often underrepresented tasks. To better understand this phenomenon, we derive an analysis to quantify this observed ``unfairness'' and draw attention to the factors that motivate such disparate speed-ups to emerge. Further, guided by these insights, the paper proposes a mitigation strategy designed to reduce speed-up disparities and validates the approach across several model pairs, revealing on average a 12% improvement in our fairness metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。