通过固定锚点分阶段学习,提升噪声下的说话人识别鲁棒性
A Stage-Wise Learning Strategy with Fixed Anchors for Robust Speaker Verification
- 分阶段训练:先建说话人边界,再用固定锚点约束噪声输入
- 在多种噪声下识别准确率显著提升,尤其在强噪声场景表现更优
- 适合需要高鲁棒性的实际语音识别系统部署
在噪声环境下学习鲁棒的说话人表示面临巨大挑战,需同时兼顾判别性与抗噪性。本文提出一种基于锚点的分阶段学习策略,首先训练基础模型以建立判别性说话人边界,随后从中提取固定锚点嵌入作为稳定参考。最后,复制基础模型,在噪声输入上进行微调,并通过强制靠近对应固定锚点嵌入,保持说话人身份不变。实验表明,该策略优于传统联合优化方法,尤其在保持判别力的同时增强抗噪能力。所提方法在多种噪声条件下均实现一致提升,可能得益于对边界稳定与差异抑制的分离处理。
原文摘要 · Abstract (English)
Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant properties. In this work, we proposed an anchor-based stage-wise learning strategy for robust speaker representation learning. Specifically, our approach begins by training a base model to establish discriminative speaker boundaries, and then extract anchor embeddings from this model as stable references. Finally, a copy of the base model is fine-tuned on noisy inputs, regularized by enforcing proximity to their corresponding fixed anchor embeddings to preserve speaker identity under distortion. Experimental results suggest that this strategy offers advantages over conventional joint optimization, particularly in maintaining discrimination while improving noise robustness. The proposed method demonstrates consistent improvements across various noise conditions, potentially due to its ability to handle boundary stabilization and variation suppression separately.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。