arXiv:2410.04785eess.AScs.SD2024-10被引 3

用类脑脉冲神经网络实现超低功耗语音增强,适合边缘设备。

Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet

  • 基于脉冲神经网络与频段融合策略,捕捉语音全局与局部特征。
  • 在Intel N-DNS挑战赛中,语音质量与能效均超越现有方法,获算法组冠军。
  • 创新脉冲神经元模型提升多尺度时序处理能力,适合听觉辅助设备等场景。

语音增强对提升各类音频设备中的语音可懂度与质量至关重要。近年来,深度学习方法显著提升了语音增强性能,但计算开销高,难以部署于耳机、助听器等大量边缘设备。本文提出基于类脑脉冲神经网络(SNN)的超低功耗语音增强系统Spiking-FullSubNet,采用全带宽与子带融合策略,有效捕获全局与局部频谱信息。为提升子带建模效率,引入受人耳外周听觉系统敏感性启发的频率分区方法。同时提出新型脉冲神经元模型,可动态调控输入信息整合与遗忘,增强SNN的多尺度时序处理能力,对语音去噪至关重要。在最新的Intel神经形态深度降噪(N-DNS)挑战赛数据集上,Spiking-FullSubNet在语音质量和能效指标上大幅领先当前最优方法,斩获算法组冠军,为边缘端超低功耗语音增强开辟新路径。源代码与模型权重已公开于https://github.com/haoxiangsnr/spiking-fullsubnet。

原文摘要 · Abstract (English)

Speech enhancement is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved speech enhancement performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultra-low-power speech enhancement system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and sub-band fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive sub-band modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multi-scale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (Algorithmic Track), opening up a myriad of opportunities for ultra-low-power speech enhancement at the edge. Our source code and model checkpoints are publicly available at https://github.com/haoxiangsnr/spiking-fullsubnet.

类脑计算语音增强脉冲神经网络边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。