arXiv:2502.05905cs.CV2025-02ICLR被引 27

轻量化脉冲神经网络,高效部署于边缘设备。

QP-SNN: Quantized and Pruned Spiking Neural Networks

  • 融合量化与结构化剪枝,降低存储与计算开销。
  • 提出权重重缩放和基于时空尖峰奇异值的剪枝准则,性能提升显著。
  • 适合资源受限场景,如边缘智能计算中的高效模型部署。

类脑脉冲神经网络(SNN)通过稀疏尖峰编码信息,以异步事件驱动方式运行,具备极高的能效优势。然而当前SNN研究主要聚焦于大规模模型性能提升,限制了其在资源受限边缘设备上的应用。本文提出一种硬件友好、轻量化的SNN——QP-SNN,旨在实现高性能SNN在资源受限场景的有效部署。首先构建集成均匀量化与结构化剪枝的基线模型(QP-SNN baseline),虽大幅降低存储与计算成本,但存在性能下降问题。为此,深入分析量化与剪枝导致性能退化的根源,提出针对性解决方案:针对权重量化,设计权重重缩放策略,更高效利用位宽以增强表征能力;针对结构化剪枝,提出基于时空尖峰活动奇异值的新剪枝准则,实现更精准地移除冗余卷积核。大量实验表明,将两项改进融入基线后,QP-SNN达到当前最优性能与效率,凸显其在边缘智能计算中部署SNN的巨大潜力。

原文摘要 · Abstract (English)

Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to encode information and operate in an asynchronous event-driven manner, offering a highly energy-efficient paradigm for machine intelligence. However, the current SNN community focuses primarily on performance improvement by developing large-scale models, which limits the applicability of SNNs in resource-limited edge devices. In this paper, we propose a hardware-friendly and lightweight SNN, aimed at effectively deploying high-performance SNN in resource-limited scenarios. Specifically, we first develop a baseline model that integrates uniform quantization and structured pruning, called QP-SNN baseline. While this baseline significantly reduces storage demands and computational costs, it suffers from performance decline. To address this, we conduct an in-depth analysis of the challenges in quantization and pruning that lead to performance degradation and propose solutions to enhance the baseline's performance. For weight quantization, we propose a weight rescaling strategy that utilizes bit width more effectively to enhance the model's representation capability. For structured pruning, we propose a novel pruning criterion using the singular value of spatiotemporal spike activities to enable more accurate removal of redundant kernels. Extensive experiments demonstrate that integrating two proposed methods into the baseline allows QP-SNN to achieve state-of-the-art performance and efficiency, underscoring its potential for enhancing SNN deployment in edge intelligence computing.

脉冲神经网络边缘计算模型压缩量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。