对比不同推测解码策略的能耗,揭示影响能效的关键因素。
Benchmarking the Energy Savings with Speculative Decoding Strategies
- 系统评估多种推测解码方法的能耗表现。
- 发现模型规模、策略类型和数据特征显著影响能效。
- 为降低大模型推理能耗提供实证依据,适合能效优化研究者。
推测解码已成为降低大语言模型推理延迟和成本的有效方法,但对其能耗的关注仍显不足。本文针对推测解码策略的能耗需求开展全面调研,深入分析模型规模与类型、推测解码策略及数据集特性等因素对能效优化的影响,为提升大模型推理能效提供实证支持。
原文摘要 · Abstract (English)
Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requirements of these models. To address this gap, this paper presents a comprehensive survey of energy requirements of speculative decoding strategies, with detailed analysis on how various factors -- model size and family, speculative decoding strategies, and dataset characteristics -- influence the energy optimizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。