用硬件加速推测性解码,提升大模型推理效率
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
- 在硬件层面支持推测性解码,优化生成过程
- 显著提升大模型推理性能与能效
- 适合需要高效推理的部署场景
大规模语言模型(LLMs)通过理解与生成类人文本,彻底改变了自然语言处理领域。然而,随着模型日益复杂,其规模带来的计算挑战也愈发严峻。本文提出硬件加速解码(HADES),一种全新的方法,旨在提升LLM的性能与能效。我们设计了一种支持硬件级推测性解码的LLM加速器,这一概念此前未见于现有文献。研究表明,推测性解码可显著提高LLM操作的效率,为更先进、实用的模型应用铺平道路。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized natural language processing by understanding and generating human-like text. However, the increasing demand for more sophisticated LLMs presents significant computational challenges due to their scale and complexity. This paper introduces Hardware Accelerated Decoding (HADES), a novel approach to enhance the performance and energy efficiency of LLMs. We address the design of an LLM accelerator with hardware-level speculative decoding support, a concept not previously explored in existing literature. Our work demonstrates how speculative decoding can significantly improve the efficiency of LLM operations, paving the way for more advanced and practical applications of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。