arXiv:2506.01166cs.ARcs.AI2025-06中稿 · publication at MOC…被引 1

通过虚拟扩展架构,让稀疏神经网络加速器更省电省面积。

VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration

  • 基于稀疏性动态虚拟扩展计算阵列,用相同硬件完成更大矩阵运算。
  • 在16纳米工艺下,相比基准架构面积节省37%,功耗效率提升68%。
  • 通用性强,适用于任意稀疏度的深度神经网络,适合边缘AI场景。

利用高度非结构化稀疏性是提升深度神经网络(DNN)加速器效率的有前景方法,尤其对新兴边缘智能应用至关重要。本文提出VUSA,一种可随当前稀疏性动态虚拟扩展的脉动阵列架构,可在不增加物理乘加单元数量的前提下执行更大规模的矩阵乘法。在商用16纳米工艺下,该架构相较基线脉动阵列,在保持峰值性能不变的情况下,实现面积节省37%、功耗效率提升68%。同时,该架构支持任意稀疏度的DNN加速,包括无稀疏性场景,具备应用无关性,适用于通用人工智能加速。

原文摘要 · Abstract (English)

Leveraging high degrees of unstructured sparsity is a promising approach to enhance the efficiency of deep neural network DNN accelerators - particularly important for emerging Edge-AI applications. We introduce VUSA, a systolic-array architecture that virtually grows based on the present sparsity to perform larger matrix multiplications with the same number of physical multiply-accumulate MAC units. The proposed architecture achieves saving by 37% and 68% in area and power efficiency, respectively, at the same peak-performance, compared to a baseline systolic array architecture in a commercial 16-nm technology. Still, the proposed architecture supports acceleration for any DNN with any sparsity - even no sparsity at all. Thus, the proposed architecture is application-independent, making it viable for general-purpose AI acceleration.

稀疏加速脉动阵列边缘AI能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。