arXiv:2606.05911cs.SDcs.LG2026-06TPAMI

混合神经网络降低语音增强计算量,兼顾性能与能效。

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

论文配图:DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement
图 1 · 摘自论文原文
  • 双分支结构融合ANN与SNN,SNN省电、ANN补信息损失。
  • 在三个数据集上性能优异,计算复杂度降低7.5倍。
  • 适合边缘设备部署,尤其关注低功耗语音处理的场景。

基于人工神经网络(ANN)的语音增强方法虽表现优秀,但高计算复杂度和高能耗限制了其在前端处理任务中的实际应用。脉冲神经网络(SNN)虽具低功耗潜力,但其离散二值激活和复杂的时空动态常导致信息损失。本文提出双分支混合神经网络(DBHN-Net),通过集成ANN与SNN分支:SNN分支降低功耗,ANN分支缓解信息丢失。设计了频带分割与时频-梅巴(TF-Mamba)模块,同步压缩能耗并提升性能;引入脉冲特征提取组(SFEG)与信息转换块(ITB),结合残差连接,进一步优化特征表示。为促进双分支信息融合,设计交互模块与时频交叉注意力融合模块(TF-Cross Attention-Fusion),实现多阶段信息交换,并自适应引导SNN保留关键信息。实验表明,该模型在三个公开数据集上保持优异性能,平均计算复杂度较基线模型降低7.5倍。

原文摘要 · Abstract (English)

Although artificial neural network (ANN) based speech enhancement (SE) methods demonstrate excellent performance, the high computational complexity and high energy consumption hinder their deployment in practical front-end processing tasks.} Currently, the spiking neural networks (SNNs) have shown potential in reducing power consumption. However, the discrete binary activation and complex spatio-temporal dynamics of SNNs often result in information loss. The current challenge therefore focuses on how to maintain performance and reduce computational complexity. To address this issue, this work propose a Dual-Branch Hybrid Neural (DBHN) Network. 1) In terms of network architecture: A dual-branch network integrating ANN and SNN was designed, where the SNN branch reduces power consumption while the ANN branch addresses information loss; The BandSplit and Time-Frequency (TF) -Mamba modules were developed to simultaneously compress energy consumption and enhance model performance; Spiking Feature Extraction Group (SFEG) and Information Transformation Block (ITB) components were implemented with residual connections to mitigate information loss while further refining feature representations. 2) To facilitate inter-branch information fusion: An Interaction module was designed to promote information exchange at various stages of the dual-branch network; A TF-Cross Attention-Fusion module was designed to perform time-frequency domain fusion of dual-branch information while data-adaptively guiding the SNN branch to retain more critical information. Results show that the proposed model maintains superior performance across three public datasets while achieving an average 7.5 fold reduction in computational complexity compared to baseline models.

语音增强神经网络低功耗混合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。