arXiv:2501.00348cs.SDcs.AI2025-01

提出时空重构与非对齐残差,提升脉冲神经网络语音分类性能

Temporal Information Reconstruction and Non-Aligned Residual in Spiking Neural Networks for Speech Classification

  • 借鉴人脑分级处理机制重构时序维度,实现多尺度时序信息学习
  • 在SSC数据集上达81.02%准确率,SHD上达96.04%,刷新SNN模型纪录
  • 适用于需要低功耗、高效率的语音识别场景,尤其适合边缘设备

当前多数脉冲神经网络(SNN)仅使用单一时间分辨率处理语音分类任务,难以捕捉输入数据在不同时间尺度上的信息。此外,许多模型中子模块前后数据时长不一致,导致有效残差连接无法应用,影响训练优化。为此,一方面,我们借鉴人脑语音处理的分层机制,提出时序重构(TR)方法,重构音频频谱的时序维度,使SNN能在不同时间分辨率下学习输入数据信息,从而建模更全面的语义特征;另一方面,通过分析音频数据特性,提出非对齐残差(NAR)方法,使残差连接可应用于时长不同的两段音频数据。我们在Spiking Speech Commands(SSC)、Spiking Heidelberg Digits(SHD)和Google Speech Commands v0.02(GSC)数据集上进行了大量实验。结果表明,在SSC上所有SNN模型中达到81.02%的测试分类准确率,为当前最优;在SHD上达到96.04%的分类准确率,同样为所有模型中的最佳表现。

原文摘要 · Abstract (English)

Recently, it can be noticed that most models based on spiking neural networks (SNNs) only use a same level temporal resolution to deal with speech classification problems, which makes these models cannot learn the information of input data at different temporal scales. Additionally, owing to the different time lengths of the data before and after the sub-modules of many models, the effective residual connections cannot be applied to optimize the training processes of these models.To solve these problems, on the one hand, we reconstruct the temporal dimension of the audio spectrum to propose a novel method named as Temporal Reconstruction (TR) by referring the hierarchical processing process of the human brain for understanding speech. Then, the reconstructed SNN model with TR can learn the information of input data at different temporal scales and model more comprehensive semantic information from audio data because it enables the networks to learn the information of input data at different temporal resolutions. On the other hand, we propose the Non-Aligned Residual (NAR) method by analyzing the audio data, which allows the residual connection can be used in two audio data with different time lengths. We have conducted plentiful experiments on the Spiking Speech Commands (SSC), the Spiking Heidelberg Digits (SHD), and the Google Speech Commands v0.02 (GSC) datasets. According to the experiment results, we have achieved the state-of-the-art (SOTA) result 81.02\% on SSC for the test classification accuracy of all SNN models, and we have obtained the SOTA result 96.04\% on SHD for the classification accuracy of all models.

脉冲神经网络语音分类时序建模残差连接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。