arXiv:2604.25624eess.AS2026-04被引 2

用双通道输入提升降噪后语音的说话人识别鲁棒性

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

论文配图:UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
图 1 · 摘自论文原文
  • 将噪声与增强语音作为双通道输入,通过U-Net融合特征
  • 在多个含噪测试集上表现优于现有方法,提升显著
  • 适合需要高鲁棒性的实际语音识别场景

在噪声环境下,联合训练语音降噪与说话人嵌入网络是常用方法。然而,该范式未能充分利用大规模语音降噪预训练带来的泛化能力,且降噪过程未显式保留说话人信息。为此,本文提出可扩展的基于U-Net的融合框架(UF-EMA),将噪声语音与增强语音作为多通道输入,使说话人编码器更有效利用说话人特征。同时,采用指数移动平均策略对在干净语音上预训练的说话人编码器进行优化,缓解过拟合,实现从干净到噪声条件的平滑过渡。在多个含噪测试集上的实验结果表明,该方法具有明显优势。

原文摘要 · Abstract (English)

The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness benefits inherent in large-scale speech enhancement pre-training. Moreover, maintaining the speaker information in the denoised speech is not an explicit objective of the speech enhancement process. To address these limitations, we proposed a scalable \textbf{U}Net-based \textbf{F}usion framework (UF-EMA) that considers the noisy and enhanced speech as a multi-channel input, thereby enabling the speaker encoder to exploit speaker information effectively. In addition, an \textbf{E}xponential \textbf{M}oving \textbf{A}verage strategy is applied to a speaker encoder pre-trained on clean speech to mitigate overfitting and facilitate a smooth transition from clean to noisy conditions. Experimental results on multiple noise-contaminated test sets showcase the superiority of the proposed approach.

说话人识别降噪U-Net鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。