提出时间正则化训练法,让脉冲神经网络更稳定、泛化更强。
Temporal Regularization Training: Unleashing the Potential of Spiking Neural Networks
- 用随时间衰减的正则化,强化早期时间步的学习约束。
- 在静态与动态数据集上均显著缓解过拟合,提升模型泛化能力。
- 适合追求低功耗、高鲁棒性的脉冲神经网络研究者。
脉冲神经网络(SNN)因其事件驱动和低功耗特性受到广泛关注,特别适用于类脑数据处理。然而,直接训练的SNN存在严重的时间梯度消失和过拟合问题,制约其性能与泛化能力。本文提出时间正则化训练(TRT)方法,通过时间衰减的正则化机制,优先对早期时间步施加更强约束,以释放SNN的潜力。理论分析表明,TRT能有效缓解时间梯度消失。在静态图像数据集和动态类脑数据集上的实验验证了其有效性:TRT显著减少过拟合,促使SNN收敛至平坦局部极小值,提升泛化性能。进一步通过追踪训练过程中SNN的费舍尔信息发现,信息逐步集中在早期时间步。TRT通过引导网络学习早期富含信息的鲁棒特征,实现了模型泛化能力的显著提升。
原文摘要 · Abstract (English)
Spiking Neural Networks (SNNs) have received widespread attention due to their event-driven and low-power characteristics, making them particularly effective for processing neuromorphic data. Recent studies have shown that directly trained SNNs suffer from severe temporal gradient vanishing and overfitting issues, which fundamentally constrain their performance and generalizability. This paper unveils a temporal regularization training (TRT) memthod, designed to unleash the generalization and performance potential of SNNs through a time-decaying regularization mechanism that prioritizes early timesteps with stronger constraints. We perform theoretical analysis to reveal TRT's ability on mitigating the temporal gradient vanishment. To validate the effectiveness of TRT, we conduct experiments on both static image datasets and dynamic neuromorphic datasets, perform analysis of their results, demonstrating that TRT can effectively mitigate overfitting and help SNNs converge into flatter local minima with better generalizability. Furthermore, we establish a theoretical interpretation of TRT's temporal regularization mechanism by analyzing the temporal information dynamics inside SNNs. We track the Fisher information of SNNs during training process, showing that Fisher information progressively concentrates in early timesteps. The time-decaying regularization mechanism implemented in TRT effectively guides the network to learn robust features in early timesteps with rich information, thereby leading to significant improvements in model generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。