提出多维分组架构,显著提升脉冲视觉变压器能效与性能。
Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
- 跨时序、空间与结构维度分组计算,降低复杂度。
- 在ImageNet等基准上实现超高能效与高精度平衡。
- 适合追求低功耗智能推理的硬件部署场景。
脉冲神经网络(SNNs)相比人工神经网络(ANNs)具有更优的能效,但在应用于脉冲视觉变压器(S-ViTs)时,训练与推理性能存在显著不足。现有方法如ANN-SNN转换和时空反向传播(STBP)存在固有局限,难以同时优化内存占用、准确率与能耗。为此,本文提出Ge²mS-T,通过在时序、空间与网络结构维度上实施分组计算,构建新型架构。具体地,提出基于分组指数编码的积分发放模型(ExpG-IF),实现无损转换且训练开销恒定,精确调控脉冲模式;设计分组脉冲自注意力机制(GW-SSA),通过多尺度标记分组与混合注意力-卷积框架内的免乘法操作降低计算复杂度。实验表明,该方法在挑战性基准上实现了卓越性能与超高的能效表现。据我们所知,这是首个系统性建立多维分组计算以解决S-ViTs中内存开销、学习能力与能耗预算三者矛盾的工作。代码已开源:https://github.com/hzc1208/Ge2mST。
原文摘要 · Abstract (English)
Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs). However, they encounter significant deficiencies in training and inference metrics when applied to Spiking Vision Transformers (S-ViTs). Existing paradigms including ANN-SNN Conversion and Spatial-Temporal Backpropagation (STBP) suffer from inherent limitations, precluding concurrent optimization of memory, accuracy and energy consumption. To address these issues, we propose Ge$^\text{2}$mS-T, a novel architecture implementing grouped computation across temporal, spatial and network structure dimensions. Specifically, we introduce the Grouped-Exponential-Coding-based IF (ExpG-IF) model, enabling lossless conversion with constant training overhead and precise regulation for spike patterns. Additionally, we develop Group-wise Spiking Self-Attention (GW-SSA) to reduce computational complexity via multi-scale token grouping and multiplication-free operations within a hybrid attention-convolution framework. Experiments confirm that our method can achieve superior performance with ultra-high energy efficiency on challenging benchmarks. To our best knowledge, this is the first work to systematically establish multi-dimensional grouped computation for resolving the triad of memory overhead, learning capability and energy budget in S-ViTs. Code is available at https://github.com/hzc1208/Ge2mST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。