用尖峰信号指导剪枝,让视觉脉冲模型更省电高效
Vision SmolMamba: Spike-Guided Token Pruning for Energy-Efficient Spiking State-Space Vision Models

- 基于尖峰激活强度和首次尖峰延迟动态剪枝冗余视觉令牌
- 相比现有脉冲模型,能耗降低至少1.5倍,精度不降反升
- 适合低功耗边缘视觉设备,尤其适用于事件相机数据
脉冲Transformer在长程视觉建模中展现出潜力,但其二次方复杂度的令牌交互与脉冲神经计算的稀疏性不匹配。为此,我们提出Vision SmolMamba,一种融合尖峰驱动机制与线性时间选择性递归的节能脉冲状态空间架构。核心是尖峰引导的时空令牌剪枝器(SST-TP),通过尖峰激活强度和首次尖峰延迟评估令牌重要性,逐步移除冗余信息,保留关键时空特征,实现高稀疏下的高效扩展。基于此,所提出的SmolMamba模块将尖峰事件直接嵌入双向状态空间递归,构建高效的长程视觉主干网络。在ImageNet-1K、CIFAR10/100、CIFAR10-DVS和DVS128 Gesture等静态与事件基准上,Vision SmolMamba始终取得更优的准确率-效率权衡。特别地,在保持或提升精度的同时,能耗比现有脉冲Transformer基线及脉冲Mamba变体降低至少1.5倍。结果表明,结合尖峰引导的令牌稀疏与状态空间建模,为脉冲视觉系统提供了一种可扩展且节能的新范式。
原文摘要 · Abstract (English)
Spiking Transformers have shown strong potential for long-range visual modeling through spike-driven self-attention. However, their quadratic token interactions remain fundamentally misaligned with the sparse and event-driven nature of spiking neural computation. To address this limitation, we propose Vision SmolMamba, an energy-efficient spiking state-space architecture that integrates spike-driven dynamics with linear-time selective recurrence. The key idea is a Spike-Guided Spatio-Temporal Token Pruner (SST-TP), which estimates token importance using both spike activation strength and first-spike latency. This mechanism progressively removes redundant tokens while preserving salient spatio-temporal information, enabling efficient scaling with token sparsity. Based on this mechanism, the proposed SmolMamba block incorporates spike events directly into bidirectional state-space recurrence, forming a spiking state-space vision backbone for efficient long-range modeling. Extensive experiments on both static and event-based benchmarks, including ImageNet-1K, CIFAR10/100, CIFAR10-DVS, and DVS128 Gesture, demonstrate that Vision SmolMamba consistently achieves superior accuracy-efficiency trade-offs. In particular, it reduces the estimated energy cost by at least 1.5x compared with prior spiking Transformer baselines and a Spiking Mamba variant while maintaining competitive or improved accuracy. These results demonstrate that combining spike-guided token sparsity with state-space modeling offers a scalable and energy-efficient paradigm for spiking vision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。