针对噪声环境下的语音验证,提出分治式专家网络框架。
Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
- 根据噪声特征自动分配输入到专用专家网络
- 在多个信噪比下均优于基线模型
- 适合实际应用中复杂噪声场景的语音识别
在噪声环境下实现鲁棒的说话人验证仍是开放挑战。传统深度学习方法通过学习统一的鲁棒说话人表示空间来应对多种背景噪声,取得了显著进步。本文提出一种噪声条件化的专家混合框架,将特征空间分解为针对特定噪声的子空间以提升说话人验证性能。具体包括:基于噪声条件的专家路由机制、基于通用模型的专家专业化策略,以及信噪比递减的课程学习协议,共同提升了模型在多样噪声条件下的鲁棒性与泛化能力。该方法可依据输入中的噪声信息自动将样本路由至对应专家网络,每个专家专注于特定噪声特性,同时保留说话人身份信息。大量实验表明,该方法在多个噪声条件下均持续优于基线模型。
原文摘要 · Abstract (English)
Robust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In contrast, this paper presents a noise-conditioned mixture-ofexperts framework that decomposes the feature space into specialized noise-aware subspaces for speaker verification. Specifically, we propose a noise-conditioned expert routing mechanism, a universal model based expert specialization strategy, and an SNR-decaying curriculum learning protocol, collectively improving model robustness and generalization under diverse noise conditions. The proposed method can automatically route inputs to expert networks based on noise information derived from the inputs, where each expert targets distinct noise characteristics while preserving speaker identity information. Comprehensive experiments demonstrate consistent superiority over baselines
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。