arXiv:2412.12858cs.LGcs.AI2024-12被引 3

用脉冲网络和课程学习蒸馏,让语音指令识别更高效节能。

Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation

  • 提出全脉冲驱动的SpikeSCR框架,结合全局-局部结构提升长期学习能力。
  • 通过课程学习蒸馏,时间步数减少60%,能耗降低54.8%仍保持高性能。
  • 适合边缘神经形态计算中需低功耗处理长序列的应用场景。

脉冲神经网络(SNNs)固有的动态特性和事件驱动机制使其在处理时间信息时具有天然优势,能利用嵌入的时间序列作为时间步。近期研究已证明SNN在语音指令识别中有效,通过使用较长的时间步实现高精度。然而,过长的时间步会增加边缘计算部署负担。为此,本文提出两个关键组件:1)提出高性能全脉冲驱动框架SpikeSCR,采用全局-局部混合结构实现高效表征学习,具备长时序学习能力;2)设计基于课程学习的知识蒸馏方法(KDCL),将简单课程中学到的有价值表示逐步迁移到复杂课程,仅带来轻微性能损失,在能效与性能间取得平衡。在三个基准数据集——Spiking Heidelberg Dataset (SHD)、Spiking Speech Commands (SSC) 和 Google Speech Commands (GSC) V2 上评估显示,SpikeSCR在相同时间步下优于现有SOTA方法。进一步应用KDCL后,时间步减少60%,能耗下降54.8%,同时性能接近最新SOTA结果。本工作为边缘神经形态系统中长序列时间处理提供了重要思路。

原文摘要 · Abstract (English)

The intrinsic dynamics and event-driven nature of spiking neural networks (SNNs) make them excel in processing temporal information by naturally utilizing embedded time sequences as time steps. Recent studies adopting this approach have demonstrated SNNs' effectiveness in speech command recognition, achieving high performance by employing large time steps for long time sequences. However, the large time steps lead to increased deployment burdens for edge computing applications. Thus, it is important to balance high performance and low energy consumption when detecting temporal patterns in edge devices. Our solution comprises two key components. 1). We propose a high-performance fully spike-driven framework termed SpikeSCR, characterized by a global-local hybrid structure for efficient representation learning, which exhibits long-term learning capabilities with extended time steps. 2). To further fully embrace low energy consumption, we propose an effective knowledge distillation method based on curriculum learning (KDCL), where valuable representations learned from the easy curriculum are progressively transferred to the hard curriculum with minor loss, striking a trade-off between power efficiency and high performance. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands (GSC) V2. Our experimental results demonstrate that SpikeSCR outperforms current state-of-the-art (SOTA) methods across these three datasets with the same time steps. Furthermore, by executing KDCL, we reduce the number of time steps by 60% and decrease energy consumption by 54.8% while maintaining comparable performance to recent SOTA results. Therefore, this work offers valuable insights for tackling temporal processing challenges with long time sequences in edge neuromorphic computing systems.

脉冲神经网络语音识别边缘计算知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。