通过虚拟扩展架构,让稀疏神经网络加速器更省电省面积。
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
- 基于稀疏性动态虚拟扩展计算阵列,用相同硬件完成更大矩阵运算。
- 在16纳米工艺下,相比基准架构面积节省37%,功耗效率提升68%。
- 通用性强,适用于任意稀疏度的深度神经网络,适合边缘AI场景。
利用高度非结构化稀疏性是提升深度神经网络(DNN)加速器效率的有前景方法,尤其对新兴边缘智能应用至关重要。本文提出VUSA,一种可随当前稀疏性动态虚拟扩展的脉动阵列架构,可在不增加物理乘加单元数量的前提下执行更大规模的矩阵乘法。在商用16纳米工艺下,该架构相较基线脉动阵列,在保持峰值性能不变的情况下,实现面积节省37%、功耗效率提升68%。同时,该架构支持任意稀疏度的DNN加速,包括无稀疏性场景,具备应用无关性,适用于通用人工智能加速。
原文摘要 · Abstract (English)
Leveraging high degrees of unstructured sparsity is a promising approach to enhance the efficiency of deep neural network DNN accelerators - particularly important for emerging Edge-AI applications. We introduce VUSA, a systolic-array architecture that virtually grows based on the present sparsity to perform larger matrix multiplications with the same number of physical multiply-accumulate MAC units. The proposed architecture achieves saving by 37% and 68% in area and power efficiency, respectively, at the same peak-performance, compared to a baseline systolic array architecture in a commercial 16-nm technology. Still, the proposed architecture supports acceleration for any DNN with any sparsity - even no sparsity at all. Thus, the proposed architecture is application-independent, making it viable for general-purpose AI acceleration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。