用仿生眼动机制提升脉冲神经网络视觉模型性能
Spiking Vision Transformer with Saccadic Attention

- 借鉴生物眼动机制,设计脉冲注意力方法增强时空交互
- 在多个任务上达到当前最优性能,计算复杂度线性增长
- 适合对功耗敏感的边缘视觉场景,兼顾效率与精度
脉冲神经网络(SNN)与视觉变压器(ViT)的结合有望在边缘视觉应用中实现高能效与高性能。然而,基于SNN的ViT与人工神经网络(ANN)仍存在显著性能差距。本文分析发现,原始自注意力机制与时空脉冲信号不匹配,导致空间相关性下降、时间交互受限。为此,受生物眼动注意机制启发,提出创新的跳跃式脉冲自注意力(SSSA)方法:空间上采用新型脉冲分布评估查询与键的相关性;时间上引入跳跃式交互模块,动态聚焦关键视觉区域,增强整体场景理解。基于此构建了基于SNN的视觉变压器(SNN-ViT)。大量实验表明,SNN-ViT在多种视觉任务中达到先进水平,计算复杂度为线性。其高效性与有效性凸显了在功耗敏感的边缘视觉场景中的潜力。
原文摘要 · Abstract (English)
The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN counterparts. Here, we first analyze why SNN-based ViTs suffer from limited performance and identify a mismatch between the vanilla self-attention mechanism and spatio-temporal spike trains. This mismatch results in degraded spatial relevance and limited temporal interactions. To address these issues, we draw inspiration from biological saccadic attention mechanisms and introduce an innovative Saccadic Spike Self-Attention (SSSA) method. Specifically, in the spatial domain, SSSA employs a novel spike distribution-based method to effectively assess the relevance between Query and Key pairs in SNN-based ViTs. Temporally, SSSA employs a saccadic interaction module that dynamically focuses on selected visual areas at each timestep and significantly enhances whole scene understanding through temporal interactions. Building on the SSSA mechanism, we develop a SNN-based Vision Transformer (SNN-ViT). Extensive experiments across various visual tasks demonstrate that SNN-ViT achieves state-of-the-art performance with linear computational complexity. The effectiveness and efficiency of the SNN-ViT highlight its potential for power-critical edge vision applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。