首个支持直接训练脉冲神经网络的多核类脑架构,能效比高达1.05TFLOPS/W。
A High Energy-Efficiency Multi-core Neuromorphic Architecture for Deep SNN Training
- 每核配备前向、反向与权重梯度引擎,实现并行化训练计算。
- 相比A100 GPU,内存访问减少55%~85%,28nm工艺下能效达1.05TFLOPS/W。
- 可在FPGA上运行20核深度SNN训练与联邦学习,适合边缘自适应场景。
边缘设备需具备动态环境下的在线学习能力,类脑计算是能效受限边缘端高效智能计算的重要方向。然而现有类脑架构无法直接基于反向传播训练脉冲神经网络(SNN)。本文提出一种多核类脑架构,每个核心集成前向传播、反向传播与权重梯度计算单元,支持引擎级与核心级并行计算。通过充分利用SNN训练中的稀疏性,优化多种数据流与稀疏计算,实现1.05TFLOPS/W@FP16@28nm的高能效,相比A100 GPU在SNN训练中减少55%~85%的DRAM访问,并成功在FPGA上完成20核深度SNN训练及5节点联邦学习。本研究首次实现支持直接SNN训练的多核类脑架构,推动类脑计算在可边缘学习应用中的落地。
原文摘要 · Abstract (English)
There is a growing necessity for edge training to adapt to dynamically changing environment. Neuromorphic computing represents a significant pathway for high-efficiency intelligent computation in energy-constrained edges, but existing neuromorphic architectures lack the ability of directly training spiking neural networks (SNNs) based on backpropagation. We develop a multi-core neuromorphic architecture with Feedforward-Propagation, Back-Propagation, and Weight-Gradient engines in each core, supporting high efficient parallel computing at both the engine and core levels. It combines various data flows and sparse computation optimization by fully leveraging the sparsity in SNN training, obtaining a high energy efficiency of 1.05TFLOPS/W@ FP16 @ 28nm, 55 ~ 85% reduction of DRAM access compared to A100 GPU in SNN trainings, and a 20-core deep SNN training and a 5-worker federated learning on FPGAs. Our study develops the first multi-core neuromorphic architecture supporting the direct SNN training, facilitating the neuromorphic computing in edge-learnable applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。