arXiv:2412.14587cs.CVcs.AI2024-12中稿 · on Association for…被引 31

提出高效脉冲视觉分割模型,显著提升低功耗神经网络的性能。

Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation

  • 针对复杂架构导致脉冲稀疏问题,优化关键模块设计。
  • 引入归一化整数脉冲神经元,解决训练不稳定难题。
  • 在多个数据集上实现领先效果,适合低功耗视觉任务应用。

脉冲神经网络(SNNs)具有低功耗优势,但在图像分割任务中表现不佳。原因在于将复杂的分割网络直接转换为脉冲版本会导致性能下降和训练不收敛。为此,我们首先识别出造成脉冲发放严重减少的架构模块,并针对性改进,提出Spike2Former架构;其次,提出归一化整数脉冲神经元,解决复杂架构SNN的训练稳定性问题。在多个语义分割数据集上创下SNN新纪录:ADE20K上提升12.7% mIoU与5.0效率,VOC2012上提升14.3% mIoU与5.2效率,CityScapes上提升9.1% mIoU与6.6效率。

原文摘要 · Abstract (English)

Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0 efficiency on ADE20K, +14.3% mIoU and 5.2 efficiency on VOC2012, and +9.1% mIoU and 6.6 efficiency on CityScapes.

脉冲神经网络图像分割低功耗计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。