提出高效低功耗的边缘计算卷积核心,提升低精度推理性能。
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
- 采用时序一元二进制乘法器阵列,兼容主流DLA数据流
- 45nm工艺下面积减少75%、功耗降低62%,支持INT4/INT8精度
- 适合资源受限的边缘AI设备,推动低功耗硬件部署
深度神经网络复杂度提升对边缘设备的资源与功耗构成挑战。现有基于一元矩阵乘法的硬件虽利用数据稀疏性和低精度提升效率,但因处理单元阵列数据流差异,难以集成至商用加速器。本文提出Tempus Core,一种可扩展的一元化卷积核心,采用时序一元二进制乘法器(tub)阵列,无缝兼容NVDLA架构,保持数据流一致性并提升硬件效率。在45nm CMOS工艺下,对于INT8精度,其处理单元(PCU)相较NVDLA的CMAC单元面积减少59.3%、功耗降低15.3%;16×16处理单元阵列下,面积与功耗分别提升75%和62%,且在INT8与INT4精度下实现等面积吞吐量5倍与4倍提升。后布局布线分析显示,16×4处理单元阵列在INT4精度下仅需0.017 mm²芯片面积,总功耗仅为6.2mW。结果表明,高效低功耗的一元硬件可无缝集成于传统DLA,为边缘AI推理提供可行路径。
原文摘要 · Abstract (English)
The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent works on unary-based matrix multiplication hardware aim to leverage data sparsity and low-precision values to enhance hardware efficiency. However, the adoption and integration of such unary hardware into commercial deep learning accelerators (DLA) remain limited due to processing element (PE) array dataflow differences. This work presents Tempus Core, a convolution core with highly scalable unary-based PE array comprising of tub (temporal-unary-binary) multipliers that seamlessly integrates with the NVDLA (NVIDIA's open-source DLA for accelerating CNNs) while maintaining dataflow compliance and boosting hardware efficiency. Analysis across various datapath granularities shows that for INT8 precision in 45nm CMOS, Tempus Core's PE cell unit (PCU) yields 59.3% and 15.3% reductions in area and power consumption, respectively, over NVDLA's CMAC unit. Considering a 16x16 PE array in Tempus Core, area and power improves by 75% and 62%, respectively, while delivering 5x and 4x iso-area throughput improvements for INT8 and INT4 precisions. Post-place and route analysis of Tempus Core's PCU shows that the 16x4 PE array for INT4 precision in 45nm CMOS requires only 0.017 mm^2 die area and consumes only 6.2mW of total power. We demonstrate that area-power efficient unary-based hardware can be seamlessly integrated into conventional DLAs, paving the path for efficient unary hardware for edge AI inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。