arXiv:2505.14719cs.CVcs.AI2025-05IJCAI被引 5

用多尺度脉冲注意力提升脉冲视觉Transformer的特征提取能力

MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion

  • 引入多尺度脉冲注意力机制增强脉冲注意力块
  • 在多个主流数据集上优于现有脉冲神经网络模型
  • 适合关注能效与高性能视觉计算的研究者

脉冲神经网络(SNN)与视觉变换器架构的结合因在低功耗和高性能计算方面的潜力而受到广泛关注。然而,基于SNN的变换器架构与基于人工神经网络(ANN)的变换器架构之间仍存在显著性能差距。尽管现有方法提出了可成功整合进SNN的脉冲自注意力机制,但其整体架构在有效提取不同图像尺度特征方面仍存在瓶颈。本文针对该问题提出MSVIT。该新型脉冲驱动的变换器架构首先采用多尺度脉冲注意力(MSSA),以增强脉冲注意力块的能力。我们在多个主流数据集上验证了该方法的有效性。实验结果表明,MSVIT在性能上超越现有基于SNN的模型,成为当前基于SNN的变换器架构中的最先进方案。代码已开源:https://github.com/Nanhu-AI-Lab/MSViT。

原文摘要 · Abstract (English)

The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has garnered significant attention due to their potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT. This novel spike-driven Transformer architecture firstly uses multi-scale spiking attention (MSSA) to enhance the capabilities of spiking attention blocks. We validate our approach across various main datasets. The experimental results show that MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among SNN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT.

脉冲神经网络视觉变换器多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。