arXiv:2409.03384cs.ARcs.AI2024-09综述被引 19

对比LLM硬件加速方案,统一技术节点使比较更公平。

Hardware Acceleration of LLMs: A comprehensive survey and comparison

  • 统一不同芯片工艺,通过外推实现公平对比。
  • 在FPGA上实测部分模型,验证性能与能效可比性。
  • 适合关注硬件加速选型的研究者和工程师。

大语言模型(LLMs)已成为自然语言处理的强大工具,其理解与生成类人文本的能力彻底改变了该领域。本文全面综述了利用硬件加速器加速变压器网络以提升大型语言模型性能的多项研究工作。调查涵盖了所提出的各类框架,并从技术、处理平台(FPGA、ASIC、存内计算、GPU)、加速比、能效、性能(GOPs)及能效(GOPs/W)等方面进行定性与定量比较。主要挑战在于各方案均基于不同工艺节点实现,难以公平比较。本文主要贡献在于将性能与能效结果外推至同一工艺节点,实现理论与实践双重公平比较;我们还在多个FPGA芯片上实现了部分LLM,用于外推并完成公正评估。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful tools for natural language processing tasks, revolutionizing the field with their ability to understand and generate human-like text. In this paper, we present a comprehensive survey of the several research efforts that have been presented for the acceleration of transformer networks for Large Language Models using hardware accelerators. The survey presents the frameworks that have been proposed and then performs a qualitative and quantitative comparison regarding the technology, the processing platform (FPGA, ASIC, In-Memory, GPU), the speedup, the energy efficiency, the performance (GOPs), and the energy efficiency (GOPs/W) of each framework. The main challenge in comparison is that every proposed scheme is implemented on a different process technology making hard a fair comparison. The main contribution of this paper is that we extrapolate the results of the performance and the energy efficiency on the same technology to make a fair comparison; one theoretical and one more practical. We implement part of the LLMs on several FPGA chips to extrapolate the results to the same process technology and then we make a fair comparison of the performance.

LLM加速FPGA能效对比硬件综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。