arXiv:2409.04976cs.ARcs.AI2024-09被引 3

通过层复用与运行时配置,显著降低边缘设备的功耗和资源占用。

HYDRA: Hybrid Data Multiplexing and Run-time Layer Configurable DNN Accelerator

  • 采用层复用架构,重用单个激活函数提升计算效率。
  • 实现超90%功耗降低,峰值性能达35.21 TOPS/W。
  • 适合资源受限的边缘AI部署,尤其在能效敏感场景。

深度神经网络(DNN)在边缘节点执行高效计算面临巨大硬件资源需求挑战。本文提出HYDRA:一种混合数据复用与运行时层可配置的DNN加速器,通过层复用方法,在单层执行中进一步重用单一激活函数,并优化融合乘加(FMA)单元。该方法以迭代模式运行,复用同一硬件并可配置地执行不同层。所提架构在功耗上降低超过90%,资源利用率优于现有工作,达到35.21 TOPSW的峰值能效。同时,其面积开销比传统设计减少(N-1)倍,显著降低带宽、激活函数及层结构所需资源。实验表明,HYDRA架构可在资源受限边缘设备上实现最优的DNN计算性能。

原文摘要 · Abstract (English)

Deep neural networks (DNNs) offer plenty of challenges in executing efficient computation at edge nodes, primarily due to the huge hardware resource demands. The article proposes HYDRA, hybrid data multiplexing, and runtime layer configurable DNN accelerators to overcome the drawbacks. The work proposes a layer-multiplexed approach, which further reuses a single activation function within the execution of a single layer with improved Fused-Multiply-Accumulate (FMA). The proposed approach works in iterative mode to reuse the same hardware and execute different layers in a configurable fashion. The proposed architectures achieve reductions over 90% of power consumption and resource utilization improvements of state-of-the-art works, with 35.21 TOPSW. The proposed architecture reduces the area overhead (N-1) times required in bandwidth, AF and layer architecture. This work shows HYDRA architecture supports optimal DNN computations while improving performance on resource-constrained edge devices.

边缘计算DNN加速能效优化硬件架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。