提升说话人识别中嵌入表示的可靠性与鲁棒性
Towards Robust Uncertainty-Aware Speaker Modeling

- 引入跨说话人与同说话人感知的不确定性软最大值
- 在跨域场景下显著提升不确定性估计的可靠性
- 适合需要高可信度说话人识别的实用系统
说话人嵌入将帧级声学特征聚合为紧凑表示,用于说话人识别。近期的不确定性感知说话人建模方法通过估计嵌入的不确定性来表征其可靠性。然而,现有方法在领域迁移下常出现不确定性估计不准和校准偏差。为此,本文从估计与适配两方面提出鲁棒不确定性建模框架:首先设计了兼顾说话人间可分性与说话人内变异性的不确定性软最大值,使不确定性估计更准确反映嵌入可靠性;其次提出不确定性校准的领域自适应(UCDA)框架,缓解因领域差异导致的不确定性校准偏差。在域内与跨域基准上的大量实验表明,该方法持续提升了不确定性可靠性与说话人识别的鲁棒性。
原文摘要 · Abstract (English)
Speaker embeddings aggregate frame-level acoustic features into compact representations for speaker recognition. Recent uncertainty-aware speaker modeling approaches further characterize the reliability of speaker embeddings by estimating their associated uncertainty. However, existing methods often suffer from inaccurate uncertainty estimation and uncertainty miscalibration under domain shifts. To address these challenges, we propose a robust uncertainty modeling framework from both estimation and adaptation perspectives. Specifically, we introduce an Inter- and Intra-Speaker-Aware Uncertainty Softmax that incorporates both inter-speaker separability and intra-speaker variability into uncertainty learning, enabling uncertainty estimates to better capture the reliability of speaker embeddings. Furthermore, we propose an Uncertainty-Calibrated Domain Adaptation (UCDA) framework to mitigate uncertainty miscalibration caused by domain mismatch. Extensive experiments on both in-domain and cross-domain benchmarks demonstrate that the proposed approach consistently improves uncertainty reliability and speaker recognition robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。