提出新型时间编码矩阵乘法架构,低功耗高效支持边缘设备低精度计算。
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
- 采用时间编码实现精确矩阵乘法,替代传统随机编码方法。
- 2位精度下仅需0.01 mm²面积和4 mW功耗,4位时为0.03 mm²和9 mW。
- 适合移动与边缘设备的持续实时感知任务,能效优于现有技术。
通用矩阵乘法(GEMM)是人工智能与深度学习中广泛使用的数据处理核心运算。随着边缘计算兴起,基于一元计算的GEMM架构逐渐流行,但多为随机性、率编码系统。本文提出一种基于时间编码的新型GEMM架构tuGEMM,可实现精确计算。设计了串行与并行两种变体,分别在面积、功耗与延迟间权衡。在45 nm CMOS工艺下,针对2位、4位及8位计算进行了后综合的功耗-性能-面积(PPA)评估。结果显示,在低精度场景下,tuGEMM相比现有最先进随机一元系统具有显著面积-功耗优势:2位精度时仅需0.01 mm²和4 mW,4位时为0.03 mm²和9 mW。该特性使其特别适用于对功耗敏感的移动与边缘设备,支持持续运行的实时传感处理。
原文摘要 · Abstract (English)
General matrix multiplication (GEMM) is a ubiquitous computing kernel/algorithm for data processing in diverse applications, including artificial intelligence (AI) and deep learning (DL). Recent shift towards edge computing has inspired GEMM architectures based on unary computing, which are predominantly stochastic and rate-coded systems. This paper proposes a novel GEMM architecture based on temporal-coding, called tuGEMM, that performs exact computation. We introduce two variants of tuGEMM, serial and parallel, with distinct area/power-latency trade-offs. Post-synthesis Power-Performance-Area (PPA) in 45 nm CMOS are reported for 2-bit, 4-bit, and 8-bit computations. The designs illustrate significant advantages in area-power efficiency over state-of-the-art stochastic unary systems especially at low precisions, e.g. incurring just 0.03 mm^2 and 9 mW for 4 bits, and 0.01 mm^2 and 4 mW for 2 bits. This makes tuGEMM ideal for power constrained mobile and edge devices performing always-on real-time sensory processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。