提出SpikeX架构,通过软硬件协同优化稀疏脉冲神经网络,显著提升能效和延迟。
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks
- 采用流水线阵列架构,针对稀疏脉冲数据优化数据流,减少内存访问。
- 在不损失精度前提下,能效延迟积降低15.1至150.87倍。
- 支持网络与硬件联合设计,适用于低功耗实时脉冲计算场景。
脉冲神经网络(SNNs)是具有生物合理性的计算模型,使用类似生物神经元的脉冲二值激活函数,擅长处理时空数据,在超低功耗和实时处理方面具有优势。尽管传统人工神经网络加速器研究众多,但针对SNN高效硬件加速器的设计仍较少关注。尤其SNN存在固有的非结构化时空脉冲稀疏性,尚未被充分挖掘以实现硬件效率提升。本文提出新型脉冲阵列加速器SpikeX,应对非结构化稀疏性挑战,并结合脉冲计算特性。通过设计高效数据流,降低昂贵的多比特权重访问开销,提升跨时空计算中的数据复用与硬件利用率,显著改善能效与推理延迟。此外,考虑到网络与硬件协同设计的重要性,我们开发了联合优化方法,支持硬件感知的SNN训练与加速器架构搜索,实现网络参数与硬件架构的联合优化。该端到端协同设计方法在保持模型精度的前提下,使能量延迟积(EDP)降低15.1x–150.87x。
原文摘要 · Abstract (English)
Spiking Neural Networks (SNNs) are promising biologically plausible models of computation which utilize a spiking binary activation function similar to that of biological neurons. SNNs are well positioned to process spatiotemporal data, and are advantageous in ultra-low power and real-time processing. Despite a large body of work on conventional artificial neural network accelerators, much less attention has been given to efficient SNN hardware accelerator design. In particular, SNNs exhibit inherent unstructured spatial and temporal firing sparsity, an opportunity yet to be fully explored for great hardware processing efficiency. In this work, we propose a novel systolic-array SNN accelerator architecture, called SpikeX, to take on the challenges and opportunities stemming from unstructured sparsity while taking into account the unique characteristics of spike-based computation. By developing an efficient dataflow targeting expensive multi-bit weight data movements, SpikeX reduces memory access and increases data sharing and hardware utilization for computations spanning across both time and space, thereby significantly improving energy efficiency and inference latency. Furthermore, recognizing the importance of SNN network and hardware co-design, we develop a co-optimization methodology facilitating not only hardware-aware SNN training but also hardware accelerator architecture search, allowing joint network weight parameter optimization and accelerator architectural reconfiguration. This end-to-end network/accelerator co-design approach offers a significant reduction of 15.1x-150.87x in energy-delay-product(EDP) without comprising model accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。