arXiv:2508.12609cs.NEcs.LG2025-08

通过自集成视角提升1比特权重脉冲神经网络训练效果

A Self-Ensemble Inspired Approach for Effective Training of Binary-Weight Spiking Neural Networks

  • 将脉冲网络训练视为带噪声的二值激活集成学习
  • 仅用2个时间步在ImageNet上达82.52%准确率
  • 适合低功耗芯片部署的高效脉冲网络设计

脉冲神经网络(SNNs)因其能效优势,是类脑硬件上的有前景方案,但其非可微的放电函数使训练困难。现有方法常采用基于时间的反向传播,并用替代梯度处理非可微性。类似地,二值神经网络(BNNs)也面临相同挑战。然而,两者间的深层联系及训练技术互鉴尚未系统研究。尤其二值权重SNN训练更为困难。本文分析反向传播过程,揭示前馈SNN训练等价于带噪声注入的二值激活网络自集成。据此提出自集成启发式训练方法(SEI-BWSNN),通过多捷径结构与知识蒸馏提升性能。在Transformer中二值化前馈层,仅用2个时间步即在ImageNet上达到82.52%准确率,验证了方法有效性与二值权重SNN的潜力。

原文摘要 · Abstract (English)

Spiking Neural Networks (SNNs) are a promising approach to low-power applications on neuromorphic hardware due to their energy efficiency. However, training SNNs is challenging because of the non-differentiable spike generation function. To address this issue, the commonly used approach is to adopt the backpropagation through time framework, while assigning the gradient of the non-differentiable function with some surrogates. Similarly, Binary Neural Networks (BNNs) also face the non-differentiability problem and rely on approximating gradients. However, the deep relationship between these two fields and how their training techniques can benefit each other has not been systematically researched. Furthermore, training binary-weight SNNs is even more difficult. In this work, we present a novel perspective on the dynamics of SNNs and their close connection to BNNs through an analysis of the backpropagation process. We demonstrate that training a feedforward SNN can be viewed as training a self-ensemble of a binary-activation neural network with noise injection. Drawing from this new understanding of SNN dynamics, we introduce the Self-Ensemble Inspired training method for (Binary-Weight) SNNs (SEI-BWSNN), which achieves high-performance results with low latency even for the case of the 1-bit weights. Specifically, we leverage a structure of multiple shortcuts and a knowledge distillation-based training technique to improve the training of (binary-weight) SNNs. Notably, by binarizing FFN layers in a Transformer architecture, our approach achieves 82.52% accuracy on ImageNet with only 2 time steps, indicating the effectiveness of our methodology and the potential of binary-weight SNNs.

脉冲神经网络二值化自集成低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。