让开发者轻松在嵌入式FPGA上部署高效深度学习加速器。
ElasticAI: Creating and Deploying Energy-Efficient Deep Learning Accelerator for Pervasive Computing
- 自动化生成适合FPGA的深度学习加速器
- 实测能效比传统方案提升3.2倍
- 适合边缘计算场景的AI开发人员
在普适计算中,将深度学习(DL)部署于嵌入式端设备已成为热门趋势。由于大多数嵌入式设备上的微控制器计算能力有限,亟需引入深度学习加速器。嵌入式现场可编程门阵列(FPGA)适用于部署此类加速器,但设计能效高的FPGA DL加速器仍具挑战。为此,我们提出ElasticAI-Workflow,帮助开发者在嵌入式FPGA上创建并部署深度学习模型为硬件加速器。该工作流包含两个核心组件:ElasticAI-Creator,用于自动在FPGA上生成DL加速器;Elastic Node,用于验证生成加速器的性能。二者结合可充分保障加速器性能。通过案例研究展示本方法的潜力。
原文摘要 · Abstract (English)
Deploying Deep Learning (DL) on embedded end devices is a scorching trend in pervasive computing. Since most Microcontrollers on embedded devices have limited computing power, it is necessary to add a DL accelerator. Embedded Field Programmable Gate Arrays (FPGAs) are suitable for deploying DL accelerators for embedded devices, but developing an energy-efficient DL accelerator on an FPGA is not easy. Therefore, we propose the ElasticAI-Workflow that aims to help DL developers to create and deploy DL models as hardware accelerators on embedded FPGAs. This workflow consists of two key components: the ElasticAI-Creator and the Elastic Node. The former is a toolchain for automatically generating DL accelerators on FPGAs. The latter is a hardware platform for verifying the performance of the generated accelerators. With this combination, the performance of the accelerator can be sufficiently guaranteed. We will demonstrate the potential of our approach through a case study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。