提出无需记忆的量化方法,让脉冲神经网络更省电、更快、性能更强。
Memory-Free and Parallel Computation for Quantized Spiking Neural Networks
- 不存膜电位却能保留历史信息,解决低精度导致的性能下降
- 训练并行、推理异步,速度和能效大幅提升
- 适合部署在资源受限的边缘设备上
量化脉冲神经网络(QSNNs)具有优异的能效,适用于资源受限的边缘设备。然而,权重和膜电位的低比特表示会导致显著性能下降。本文首次发现其根本原因在于量化膜电位导致历史信息丢失。为此,提出一种无需存储膜电位的历史信息捕捉方法,提升性能的同时降低内存需求。为进一步提高计算效率,设计了并行训练与异步推理框架,显著加快训练速度与能效。结合上述方法,构建高性能高效量化脉冲神经网络MFP-QSNN。大量实验表明,该模型在多种静态与类脑图像数据集上达到当前最优性能,内存占用更低,训练速度更快。其高效性与有效性凸显其在节能类脑计算中的应用潜力。
原文摘要 · Abstract (English)
Quantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause for this decline: the loss of historical information due to the quantized membrane potential. To tackle this issue, we introduce a memory-free quantization method that captures all historical information without directly storing membrane potentials, resulting in better performance with less memory requirements. To further improve the computational efficiency, we propose a parallel training and asynchronous inference framework that greatly increases training speed and energy efficiency. We combine the proposed memory-free quantization and parallel computation methods to develop a high-performance and efficient QSNN, named MFP-QSNN. Extensive experiments show that our MFP-QSNN achieves state-of-the-art performance on various static and neuromorphic image datasets, requiring less memory and faster training speeds. The efficiency and efficacy of the MFP-QSNN highlight its potential for energy-efficient neuromorphic computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。