提出一种高效脉冲神经网络学习硬件,显著降低能耗与资源占用。
ITP-STDP: A Hardware-Efficient Intrinsic-Timing Power-of-Two Synaptic Learning Engine for On-Chip SNNs

- 通过算法与硬件协同优化,实现幂次时序脉冲学习。
- 在FPGA上能效提升4.5至219.8倍,面积仅需1.2%~3.3%。
- 适合低功耗边缘设备部署,尤其适合嵌入式SNN应用。
脉冲神经网络(SNN)有望成为第三代神经网络,在诸多应用中备受关注。然而,其庞大的突触连接导致片上学习算法训练时权重更新计算密集,造成高硬件资源消耗与能量开销。其中,脉冲时序依赖可塑性(STDP)是最广泛研究和采用的学习机制。为缓解训练带来的硬件与能耗负担,本文提出内在时序幂次二进制STDP(ITP-STDP)及其对应的原型学习引擎硬件架构。通过专用均场突触漂移模型进行动态分析,并在不同规模的SNN网络与数据集上验证。该设计在ASIC与FPGA平台上实现,相较现有方法(包括原始STDP及复杂变体)展现出更高能效、更快运行速度与更低资源占用。结果表明,该设计通过算法与硬件级优化,大幅消除了STDP的计算开销。在FPGA平台,能效提升达4.5×至219.8×;在ASIC平台,速度提升4.8×至22.01×,面积仅占先前工作的1.2%至3.3%。
原文摘要 · Abstract (English)
Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across a wide range of applications. However, the large number of synaptic connections in SNNs leads to intensive weight-update computation by on-chip learning algorithms during training, resulting in substantial hardware resource utilization and energy consumption. Among existing SNN learning algorithms, spike-timing-dependent plasticity (STDP) is one of the most extensively studied and widely adopted, serving as a fundamental learning component in SNNs. To address the hardware and energy overheads associated with SNN training, this paper presents intrinsic-timing power-of-two STDP (ITP-STDP) and its corresponding prototype learning engine hardware architecture. The proposed design is evaluated through a dedicated mean-field synaptic drift model for dynamical analysis and further validated across SNN networks of different scales and datasets. It is further implemented on both ASIC and FPGA platforms and compared with state-of-the-art approaches, including the original STDP and more complex STDP variants. The results demonstrate superior energy efficiency, higher operating speed, and substantially lower hardware resource utilization, as the proposed design eliminates most of the computational overhead of STDP through both algorithmic and hardware-level optimizations. On the FPGA platform, the proposed design improves energy efficiency by 4.5$\times$ to 219.8$\times$ over the compared designs. On the ASIC platform, the proposed design achieves a 4.8$\times$ to 22.01$\times$ speedup while consuming only 1.2% to 3.3% of the area required by prior works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。