arXiv:2506.01566cs.PFcs.AI2025-06中稿 · Version for: SAMOS…

FlexiSAGA加速器高效处理稀疏与密集矩阵运算,提升边缘设备推理速度。

FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing

  • 支持七种稀疏与密集数据流,架构可配置,适配多种DNN计算模式。
  • 在全模型稀疏-密集推理中速度提升1.41至4.28倍,优于商用和文献平台。
  • 配套专用剪枝方法,实现软硬件协同设计,适合边缘AI部署场景。

人工智能算法如深度神经网络(DNN)已广泛应用于计算机视觉与自然语言处理等领域。然而,DNN推理的高计算复杂性对资源受限的边缘设备构成挑战。利用DNN权重中的稀疏性是缓解该问题的有效途径。本文提出FlexiSAGA,一种可架构配置、数据流灵活的通用矩阵乘法(GEMM)AI硬件加速器,支持七种不同的稀疏与密集数据流,可高效处理资源密集型的DNN算子。此外,我们提出一种针对FlexiSAGA架构定制的DNN剪枝方法,使密集与稀疏卷积及全连接算子达到近最优处理效率,推动了DNN与硬件协同设计流程。实验结果表明,全模型稀疏-密集推理速度提升范围为1.41至4.28倍,显著优于商用及文献报道的加速平台。

原文摘要 · Abstract (English)

Artificial Intelligence (AI) algorithms, such as Deep Neural Networks (DNNs), have become an important tool for a wide range of applications, from computer vision to natural language processing. However, the computational complexity of DNN inference poses a significant challenge, particularly for processing on resource-constrained edge devices. One promising approach to address this challenge is the exploitation of sparsity in DNN operator weights. In this work, we present FlexiSAGA, an architecturally configurable and dataflow-flexible AI hardware accelerator for the sparse and dense processing of general matrix multiplications (GEMMs). FlexiSAGA supports seven different sparse and dense dataflows, enabling efficient processing of resource intensive DNN operators. Additionally, we propose a DNN pruning method specifically tailored towards the FlexiSAGA architecture, allowing for near-optimal processing of dense and sparse convolution and fully-connected operators, facilitating a DNN/HW co-design flow. Our results show a whole DNN sparse-over-dense inference speedup ranging from 1.41 up to 4.28, outperforming commercial and literature-reported accelerator platforms.

硬件加速稀疏计算矩阵乘法边缘AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。