提出节能型矩阵乘法阵列,兼顾精度与能效。
Energy Efficient Exact and Approximate Systolic Array Architecture for Matrix Multiplication
- 用正负部分积单元设计新型计算单元
- 相比现有设计能效提升22%~32%
- 适合容错图像处理等边缘场景
深度神经网络需要高效的矩阵乘法引擎。本文提出一种新型脉动阵列架构,采用创新的精确与近似计算单元(PE),基于能量高效的正部分积单元(PPC)和负部分积单元(NPPC)。所设计的8位精确与近似PE被应用于8×8脉动阵列,在能效上分别比现有设计降低22%和32%。为验证有效性,该PE集成至脉动阵列用于离散余弦变换(DCT)计算,输出质量优异,峰值信噪比(PSNR)达38.21 dB。在边缘检测卷积应用中,近似PE实现30.45 dB的PSNR。结果表明,该设计在保持良好输出质量的同时显著提升能效,适用于容错性强的图像与视觉处理任务。
原文摘要 · Abstract (English)
Deep Neural Networks (DNNs) require highly efficient matrix multiplication engines for complex computations. This paper presents a systolic array architecture incorporating novel exact and approximate processing elements (PEs), designed using energy-efficient positive partial product and negative partial product cells, termed as PPC and NPPC, respectively. The proposed 8-bit exact and approximate PE designs are employed in a 8x8 systolic array, which achieves a energy savings of 22% and 32%, respectively, compared to the existing design. To demonstrate their effectiveness, the proposed PEs are integrated into a systolic array (SA) for Discrete Cosine Transform (DCT) computation, achieving high output quality with a PSNR of 38.21,dB. Furthermore, in an edge detection application using convolution, the approximate PE achieves a PSNR of 30.45,dB. These results highlight the potential of the proposed design to deliver significant energy efficiency while maintaining competitive output quality, making it well-suited for error-resilient image and vision processing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。