融合噪声与增强语音,提升极端降噪环境下的说话人验证准确率
A Framework for Robust Speaker Verification in Highly Noisy Environments Leveraging Both Noisy and Enhanced Audio
- 采用孪生网络同时提取噪声和增强语音的说话人特征
- 在极低信噪比下验证准确率显著优于传统方法
- 框架轻量且兼容多种主流语音增强与验证模型
近期说话人验证技术进展虽有潜力,但在复杂声学环境下性能常大幅下降。尽管语音增强方法可提升听觉质量,但可能无意中扭曲说话人特有信息,影响验证准确性。这一问题在生成式深度神经网络(DNN)用于语音增强时尤为突出:这些网络可在极低信噪比(SNR)条件下生成可懂语音,却可能严重改变说话人独特特征。为此,我们提出一种新型神经网络框架,通过孪生结构有效融合从噪声语音和增强语音中提取的说话人嵌入。该架构充分利用两者互补信息,在严苛噪声条件下显著提升说话人验证鲁棒性。框架轻量且不依赖特定验证或增强技术,可无缝集成多种前沿方案而无需修改。实验结果表明,所提框架性能显著优于现有方法。
原文摘要 · Abstract (English)
Recent advancements in speaker verification techniques show promise, but their performance often deteriorates significantly in challenging acoustic environments. Although speech enhancement methods can improve perceived audio quality, they may unintentionally distort speaker-specific information, which can affect verification accuracy. This problem has become more noticeable with the increasing use of generative deep neural networks (DNNs) for speech enhancement. While these networks can produce intelligible speech even in conditions of very low signal-to-noise ratio (SNR), they may also severely alter distinctive speaker characteristics. To tackle this issue, we propose a novel neural network framework that effectively combines speaker embeddings extracted from both noisy and enhanced speech using a Siamese architecture. This architecture allows us to leverage complementary information from both sources, enhancing the robustness of speaker verification under severe noise conditions. Our framework is lightweight and agnostic to specific speaker verification and speech enhancement techniques, enabling the use of a wide range of state-of-the-art solutions without modification. Experimental results demonstrate the superior performance of our proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。