可参数化卷积加速器,兼顾嵌入式场景的性能与功耗
A Parameterizable Convolution Accelerator for Embedded Deep Learning Applications
- 用高层次综合工具实现可调参数的卷积加速器设计
- 在多约束下比传统设计更优,支持多种深度学习应用扩展
- 适合资源受限的嵌入式AI部署,如边缘设备
在现场可编程门阵列(FPGA)上实现的卷积神经网络(CNN)加速器通常以最大化性能(以GOPS为单位)为目标。然而,真实嵌入式深度学习(DL)应用面临延迟、功耗、面积和成本等多重约束。本文提出一种软硬件协同设计方法,利用高层次综合(HLS)工具描述CNN加速器,实现设计参数化,从而在多个设计约束下进行更有效的优化。实验结果表明,所提方法优于非参数化设计,且易于扩展至其他类型的DL应用。
原文摘要 · Abstract (English)
Convolutional neural network (CNN) accelerators implemented on Field-Programmable Gate Arrays (FPGAs) are typically designed with a primary focus on maximizing performance, often measured in giga-operations per second (GOPS). However, real-life embedded deep learning (DL) applications impose multiple constraints related to latency, power consumption, area, and cost. This work presents a hardware-software (HW/SW) co-design methodology in which a CNN accelerator is described using high-level synthesis (HLS) tools that ease the parameterization of the design, facilitating more effective optimizations across multiple design constraints. Our experimental results demonstrate that the proposed design methodology is able to outperform non-parameterized design approaches, and it can be easily extended to other types of DL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。