用脉冲时间差编码实现低功耗关键词识别,效率远超传统模型。
Towards efficient keyword spotting using spike-based time difference encoders
- 通过脉冲时间差编码将语音频谱转为脉冲信号,捕捉时间特征
- 前馈TDE网络准确率达89%,比同类模型少92%突触操作
- 结果可解释性强,适合边缘设备上高效部署
在边缘设备上实现关键词识别日益重要,但受制于极低功耗限制。本文研究了时间差编码器(TDE)在关键词识别中的性能。该神经元模型通过瞬时频率和脉冲数量编码时间差,适配类脑处理器。使用包含发音数字的TIdigits数据集,经共振峰分解与基于速率的脉冲编码。比较三种脉冲神经网络(SNN)架构:前馈TDE、前馈电流基漏电积分-发放(CuBa-LIF)及递归CuBa-LIF。结果显示,频率转换后的语音脉冲序列在时间域蕴含丰富信息,凸显时间编码的重要性。在相同突触权重下,前馈TDE网络准确率达89%,高于前馈CuBa-LIF的71%,接近递归CuBa-LIF的91%。但前馈TDE仅需递归模型9%的突触操作数。且其结果与语音关键词的频率与时间尺度特征高度相关。表明TDE是可扩展事件驱动处理时空模式的有力候选。
原文摘要 · Abstract (English)
Keyword spotting in edge devices is becoming increasingly important as voice-activated assistants are widely used. However, its deployment is often limited by the extreme low-power constraints of the target embedded systems. Here, we explore the Temporal Difference Encoder (TDE) performance in keyword spotting. This recent neuron model encodes the time difference in instantaneous frequency and spike count to perform efficient keyword spotting with neuromorphic processors. We use the TIdigits dataset of spoken digits with a formant decomposition and rate-based encoding into spikes. We compare three Spiking Neural Networks (SNNs) architectures to learn and classify spatio-temporal signals. The proposed SNN architectures are made of three layers with variation in its hidden layer composed of either (1) feedforward TDE, (2) feedforward Current-Based Leaky Integrate-and-Fire (CuBa-LIF), or (3) recurrent CuBa-LIF neurons. We first show that the spike trains of the frequency-converted spoken digits have a large amount of information in the temporal domain, reinforcing the importance of better exploiting temporal encoding for such a task. We then train the three SNNs with the same number of synaptic weights to quantify and compare their performance based on the accuracy and synaptic operations. The resulting accuracy of the feedforward TDE network (89%) is higher than the feedforward CuBa-LIF network (71%) and close to the recurrent CuBa-LIF network (91%). However, the feedforward TDE-based network performs 92% fewer synaptic operations than the recurrent CuBa-LIF network with the same amount of synapses. In addition, the results of the TDE network are highly interpretable and correlated with the frequency and timescale features of the spoken keywords in the dataset. Our findings suggest that the TDE is a promising neuron model for scalable event-driven processing of spatio-temporal patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。