arXiv:2604.23165cs.CV2026-04中稿 · KDD

用脉冲爆发增强视觉表征,提升效率与精度。

BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

论文配图:BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning
图 1 · 摘自论文原文
  • 采用双通道脉冲爆发自注意力,提升信息编码能力。
  • 在静态与事件视觉任务中均优于现有脉冲Transformer。
  • 适合低功耗神经形态计算,兼顾精度与能效。

脉冲视觉变压器(S-ViTs)为节能视觉学习提供了有前景的框架。然而,现有设计受限于二进制脉冲编码的信息容量有限,以及全局自注意力带来的密集令牌交互。本文提出BSViT,一种基于脉冲爆发驱动的视觉变压器,其核心是双通道脉冲爆发自注意力(DBSSA)机制。该机制用二进制脉冲编码查询,用脉冲爆发编码键,以增强表征能力;值路径采用双兴奋与抑制二进制通道,实现带符号调制和更丰富的脉冲交互。重要的是,整个注意力运算仅保留加法计算,兼容低功耗神经形态硬件。为进一步降低脉冲活动并引入空间先验,提出补丁邻接掩码策略,将注意力限制在局部邻域,实现结构感知稀疏性与计算开销减少。此外,脉冲爆发编码在整个网络中系统集成,使脉冲层级表征能力超越传统二进制脉冲。在静态与事件视觉基准上的大量实验表明,BSViT在准确率上持续优于现有脉冲变压器,同时保持竞争力的能效表现。

原文摘要 · Abstract (English)

Spiking Vision Transformers (S-ViTs) offer a promising framework for energy-efficient visual learning. However, existing designs remain limited by two fundamental issues: the restricted information capacity of binary spike coding and the dense token interactions introduced by global self-attention. To address these challenges, this work proposes BSViT, a burst spiking-driven Vision Transformer featuring a Dual-Channel Burst Spiking Self-Attention (DBSSA) mechanism. DBSSA encodes queries with binary spikes and keys with burst spikes to enhance representational capacity. The value pathway adopts dual excitatory and inhibitory binary channels, enabling signed modulation and richer spike interactions. Importantly, the entire attention operation preserves addition-only computation, ensuring compatibility with energy-efficient neuromorphic hardware. To further reduce spike activity and incorporate spatial priors, a patch adjacency masking strategy is introduced to restrict attention to local neighborhoods, resulting in structure-aware sparsity and reduced computational overhead. In addition, burst spike coding is systematically integrated across the network to increase spike-level representational capacity beyond conventional binary spiking. Extensive experiments on both static and event-based vision benchmarks demonstrate that BSViT consistently outperforms existing spiking Transformers in accuracy while maintaining competitive energy efficiency.

脉冲神经网络视觉表征低功耗计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。