arXiv:2603.18039cs.NEcs.LG2026-03被引 1

用平滑优化提升脉冲神经网络训练,显著缩小精度差距。

Sharpness Aware Surrogate Training for Spiking Neural Networks

  • 将SAM思想融入代理前向模型,获得精确梯度。
  • N-MNIST上硬脉冲准确率从65.7%提至94.7%。
  • 适合追求高精度脉冲网络部署的研究者。

脉冲神经网络(SNN)训练依赖代理梯度,但传统方法将非光滑前向模型与有偏梯度估计耦合。本文提出锐度感知代理训练(SAST),在基于反向传播的代理前向SNN上应用锐度感知最小化(SAM)。该方法使优化目标变为平滑经验风险,从而获得对辅助模型的精确梯度。在显式有界性和收缩假设下,推导出状态稳定性与输入利普希茨界,建立代理目标的光滑性,给出一阶SAM近似界,并证明随机SAST在独立第二小批量下的非凸收敛性。此外,分离提出局部机制命题:在局部雅可比条件下,样本参数梯度控制可导致更小的输入梯度范数。实验评估了清洁准确率、硬脉冲迁移、抗扰动鲁棒性及训练开销,均在N-MNIST和DVS Gesture数据集上进行。最明显效果是迁移差距缩小:在N-MNIST上,硬脉冲准确率从65.7%升至94.7%(最佳ρ=0.30),代理前向准确率保持高位;在DVS Gesture上,硬脉冲准确率从31.8%提升至63.3%(最佳ρ=0.40)。还明确了计算匹配、校准与理论对齐的必要控制项以实现最终实用评估。

原文摘要 · Abstract (English)

Surrogate gradients are a standard tool for training spiking neural networks (SNNs), but conventional hard forward or surrogate backward training couples a nonsmooth forward model with a biased gradient estimator. We study sharpness aware Surrogate Training (SAST), which applies sharpness aware Minimization (SAM) to a surrogate forward SNN trained by backpropagation. In this formulation, the optimization target is an ordinary smooth empirical risk, so the training gradient is exact for the auxiliary model being optimized. Under explicit boundedness and contraction assumptions, we derive compact state stability and input Lipschitz bounds, establish smoothness of the surrogate objective, provide a first order SAM approximation bound, and prove a nonconvex convergence guarantee for stochastic SAST with an independent second minibatch. We also isolate a local mechanism proposition, stated separately from the unconditional guarantees, that links per sample parameter gradient control to smaller input gradient norms under local Jacobian conditioning. Empirically, we evaluate clean accuracy, hard spike transfer, corruption robustness, and training overhead on N-MNIST and DVS Gesture. The clearest practical effect is transfer gap reduction: on N-MNIST, hard spike accuracy rises from 65.7% to 94.7% (best at $ρ=0.30$) while surrogate forward accuracy remains high; on DVS Gesture, hard spike accuracy improves from 31.8% to 63.3% (best at $ρ=0.40$). We additionally specify the compute matched, calibration, and theory alignment controls required for a final practical assessment.

脉冲神经网络代理梯度锐度感知训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。