arXiv:2606.01135cs.NEcs.SD2026-06中稿 · IJCNN2026

将脉冲神经网络引入语音识别模型,显著降低计算开销。

Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition

论文配图:Spiking and Event-driven Neuromorphic Mamba Models for Efficient Speech Recognition
图 1 · 摘自论文原文
  • 用事件驱动与脉冲机制改造SpeechMamba模型,实现高稀疏激活
  • 在LibriSpeech上达60%以上稀疏度,精度损失小于1%
  • 适合边缘设备部署,支持软硬件协同优化

深度学习极大推动了自动语音识别(ASR)的发展,使其广泛应用于智能手机、智能家居等边缘设备。然而,深度神经网络的计算与能耗需求对资源受限设备构成挑战,导致延迟高、难以实现实时交互。类脑计算通过脉冲神经网络(SNNs)和事件驱动神经网络引入激活稀疏性,将密集运算转化为稀疏计算,提供潜在解决方案。但针对ASR的类脑策略硬件效益评估仍不足。本文探索将脉冲与事件驱动神经网络应用于前沿SpeechMamba模型,提升其激活稀疏性。提出基于FATReLU激活的事件驱动SpeechMamba,在LibriSpeech上实现超过60%的激活稀疏度,精度下降低于1%。同时设计脉冲SpeechMamba,稀疏度超70%,参数量比同类SNN少30%。最后构建周期精确的事件驱动仿真器,支持算法-硬件协同探索,识别计算瓶颈并带来额外10%以上的效率提升。

原文摘要 · Abstract (English)

Deep learning has greatly advanced automatic speech recognition (ASR), enabling widespread deployment on edge devices such as smartphones and smart home systems. However, the computational and energy demands of deep neural networks pose significant challenges for such resource-constrained deployments, introducing latency and limiting real-time interaction. Neuromorphic computing offers a promising solution by introducing activation sparsity through spiking neural networks (SNNs) and event-driven neural networks, converting dense operations into sparse computations. However, a study that evaluates the hardware benefits of different neuromorphic strategies remains lacking for ASR. This paper explores spiking and event-driven neuromorphic neural networks to improve activation sparsity in the state-of-the-art SpeechMamba model for ASR. We introduce an event-driven SpeechMamba with FATReLU activation, achieving over 60% activation sparsity with less than 1% accuracy degradation on LibriSpeech. We also propose a spiking SpeechMamba that attains over 70% sparsity while using 30% fewer parameters than comparable SNNs. Finally, we develop a cycle-accurate event-driven simulator enabling flexible algorithm-hardware co-exploration, which helps us identify computational bottlenecks and yields over 10% additional efficiency improvements.

语音识别脉冲神经网络稀疏计算边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。