用多视角脉冲注意力提升语音命令识别的效率与精度
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
- 设计多视角脉冲时序自注意力模块,捕捉语音的多维时间依赖
- 在三个数据集上以更少参数实现优于现有脉冲神经网络的性能
- 适合低功耗嵌入式语音识别场景,尤其关注能效比的开发者
脉冲神经网络(SNN)通过事件驱动处理机制,为低功耗语音命令识别(SCR)提供了新路径。然而,现有基于SNN的SCR方法在建模语音的丰富时序依赖和上下文信息方面存在不足,主要受限于有限的时序建模能力和二值脉冲表示。为此,我们首先提出多视角脉冲时序感知自注意力(MSTASA)模块,结合有效的脉冲时序感知注意力与多视角学习框架,以建模语音命令中的互补时序依赖。在此基础上,我们进一步构建了完全脉冲驱动的Transformer架构SpikCommander,其融合MSTASA与脉冲上下文优化通道MLP(SCR-MLP),协同增强时序上下文建模与通道特征融合。我们在三个基准数据集——斯图加特脉冲语料库(SHD)、脉冲语音命令数据集(SSC)和谷歌语音命令数据集V2(GSC)上进行了评估。大量实验表明,SpikCommander在相近时间步数下,以更少参数持续超越现有SOTA SNN方法,验证了其在鲁棒语音命令识别中的高效性与有效性。
原文摘要 · Abstract (English)
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。