提出并行噪声分离与语音增强模型,提升嘈杂环境下的说话人识别准确率。
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
- 双U-Net结构并行学习噪声提取与语音增强,显式建模噪声特征。
- 在多个噪声场景下,等错误率(EER)降低8.4%,优于现有联合模型。
- 适合需要高鲁棒性的实际语音识别系统,如智能客服、安防监控。
噪声鲁棒的说话人验证通过联合学习语音增强(SE)与说话人验证(SV)来提升性能。然而,现有方法依赖隐式噪声抑制,训练中未显式区分噪声与语音,导致噪声分离效果不佳。尽管集成SE与SV有一定帮助,仍难以有效应对复杂噪声。近期研究表明,显式建模噪声比单纯抑制更有利于提升抗噪能力。为此,我们提出ParaNoise-SV,采用双U-Net结构,包含噪声提取(NE)网络与语音增强(SE)网络。其中,NE U-Net显式建模噪声,而SE U-Net通过并行连接接收来自NE的指导,保留说话人相关特征。实验表明,ParaNoise-SV相较于先前联合SE-SV模型,等错误率(EER)相对降低8.4%。
原文摘要 · Abstract (English)
Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on implicit noise suppression, which struggles to separate noise from speaker characteristics as they do not explicitly distinguish noise from speech during training. Although integrating SE and SV helps, it remains limited in handling noise effectively. Meanwhile, recent SE studies suggest that explicitly modeling noise, rather than merely suppressing it, enhances noise resilience. Reflecting this, we propose ParaNoise-SV, with dual U-Nets combining a noise extraction (NE) network and a speech enhancement (SE) network. The NE U-Net explicitly models noise, while the SE U-Net refines speech with guidance from NE through parallel connections, preserving speaker-relevant features. Experimental results show that ParaNoise-SV achieves a relatively 8.4% lower equal error rate (EER) than previous joint SE-SV models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。