首个纯脉冲驱动的骨架动作识别模型,大幅降低能耗。
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
- 用多流解剖脉冲嵌入生成稀疏事件流,保留骨骼数据特性。
- 引入横向脉冲拓扑路由与脉冲状态空间引擎,实现长时序建模。
- 在多个数据集上达到顶尖精度,适合边缘设备部署。
基于骨架的动作识别对多媒体应用至关重要,但传统人工神经网络(ANNs)功耗高,限制了其在资源受限边缘设备上的部署。脉冲神经网络(SNNs)提供了更节能的替代方案,但现有针对骨架数据的脉冲模型常因密集矩阵聚合、复杂多模态融合模块或非稀疏频域变换而破坏了SNN的内在稀疏性,且严重受制于脉冲神经元的短期遗忘问题。本文提出首个纯粹脉冲驱动的Transformer架构——脉冲状态空间拓扑变压器(S3T-Former),专为高效能骨架动作识别设计。我们构建了多流解剖脉冲嵌入(M-ASE),作为广义运动学微分算子,将多模态骨架特征转化为异构稀疏事件流;引入横向脉冲拓扑路由(LSTR),实现按需条件脉冲传播,确保拓扑与时间稀疏性;并设计脉冲状态空间(S3)引擎,系统捕捉长程时序动态,避免非稀疏频域处理。大量实验表明,S3T-Former在多个大规模数据集上实现具有竞争力的准确率,同时理论上显著降低能耗,建立了神经形态动作识别的新基准。
原文摘要 · Abstract (English)
Skeleton-based action recognition is crucial for multimedia applications but heavily relies on power-hungry Artificial Neural Networks (ANNs), limiting their deployment on resource-constrained edge devices. Spiking Neural Networks (SNNs) provide an energy-efficient alternative; however, existing spiking models for skeleton data often compromise the intrinsic sparsity of SNNs by resorting to dense matrix aggregations, heavy multimodal fusion modules, or non-sparse frequency domain transformations. Furthermore, they severely suffer from the short-term amnesia of spiking neurons. In this paper, we propose the Spiking State-Space Topology Transformer (S3T-Former), which, to the best of our knowledge, is the first purely spike-driven Transformer architecture specifically designed for energy-efficient skeleton action recognition. Rather than relying on heavy fusion overhead, we formulate a Multi-Stream Anatomical Spiking Embedding (M-ASE) that acts as a generalized kinematic differential operator, elegantly transforming multimodal skeleton features into heterogeneous, highly sparse event streams. To achieve true topological and temporal sparsity, we introduce Lateral Spiking Topology Routing (LSTR) for on-demand conditional spike propagation, and a Spiking State-Space (S3) Engine to systematically capture long-range temporal dynamics without non-sparse spectral workarounds. Extensive experiments on multiple large-scale datasets demonstrate that S3T-Former achieves highly competitive accuracy while theoretically reducing energy consumption compared to classic ANNs, establishing a new state-of-the-art for energy-efficient neuromorphic action recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。