用自监督特征增强语音降噪,提升清晰度与抗干扰能力。
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
- 双路径结构融合幅度相位信息,结合自监督学习提取特征。
- 在VoiceBank+DEMAND和WHAMR!数据集上优于现有方法。
- 适合语音增强、自监督学习研究者参考使用。
语音自监督学习(SSL)在各类语音处理任务中取得显著进展,但在语音增强(SE)方面仍有提升空间。本文提出BSP-MPNet,一种结合自监督特征与幅度-相位信息的双路径框架。首先通过感知对比拉伸(PCS)算法增强幅度-相位谱;随后,一个二维粗粒度编码器(MP-2DC)从增强谱中提取粗粒度特征;接着,特征分离自监督学习(FS-SSL)模型分别对幅度与相位成分生成自监督嵌入,并融合形成跨域特征表示;最后,两个并行的RNN增强多注意力(REMA)掩码解码器进一步优化特征,应用掩码并重建语音信号。我们在VoiceBank+DEMAND与WHAMR!数据集上评估BSP-MPNet,实验结果表明其在多种噪声条件下均优于现有方法,为自监督语音增强研究提供了新方向。代码已公开于GitHub。
原文摘要 · Abstract (English)
Speech self-supervised learning (SSL) has made great progress in various speech processing tasks, but there is still room for improvement in speech enhancement (SE). This paper presents BSP-MPNet, a dual-path framework that combines self-supervised features with magnitude-phase information for SE. The approach starts by applying the perceptual contrast stretching (PCS) algorithm to enhance the magnitude-phase spectrum. A magnitude-phase 2D coarse (MP-2DC) encoder then extracts coarse features from the enhanced spectrum. Next, a feature-separating self-supervised learning (FS-SSL) model generates self-supervised embeddings for the magnitude and phase components separately. These embeddings are fused to create cross-domain feature representations. Finally, two parallel RNN-enhanced multi-attention (REMA) mask decoders refine the features, apply them to the mask, and reconstruct the speech signal. We evaluate BSP-MPNet on the VoiceBank+DEMAND and WHAMR! datasets. Experimental results show that BSP-MPNet outperforms existing methods under various noise conditions, providing new directions for self-supervised speech enhancement research. The implementation of the BSP-MPNet code is available online\footnote[2]{https://github.com/AlimMat/BSP-MPNet. \label{s1}}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。