arXiv:2509.05993cs.SDeess.AS2025-09被引 3

提升说话人嵌入的鲁棒性,通过显式不确定性监督增强帧重要性评估。

Xi+: Uncertainty Supervision for Robust Speaker Embedding

  • 引入时序注意力捕捉帧间上下文关系,更准确估计每帧可靠性
  • 设计新型随机方差损失函数,显式指导不确定性学习
  • 在VoxCeleb1-O和NIST SRE 2024上分别提升约10%和11%

说话人识别系统性能受情绪、语言等多重因素影响,单个语音帧对整体表示的贡献不均,因此需估计各帧的重要性或可靠性。xi-vector模型通过不确定性估计为不同帧分配权重,但其不确定性仅通过分类损失隐式训练,未考虑帧间时序关系,可能导致监督不足。本文提出改进架构xi+,引入时序注意力模块,实现上下文感知的帧级不确定性建模;同时设计新型损失函数——随机方差损失(Stochastic Variance Loss),显式监督不确定性学习。实验表明,在VoxCeleb1-O测试集上性能提升约10%,在NIST SRE 2024评估集上提升约11%。

原文摘要 · Abstract (English)

There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Since individual speech frames do not contribute equally to the utterance-level representation, it is essential to estimate the importance or reliability of each frame. The xi-vector model addresses this by assigning different weights to frames based on uncertainty estimation. However, its uncertainty estimation model is implicitly trained through classification loss alone and does not consider the temporal relationships between frames, which may lead to suboptimal supervision. In this paper, we propose an improved architecture, xi+. Compared to xi-vector, xi+ incorporates a temporal attention module to capture frame-level uncertainty in a context-aware manner. In addition, we introduce a novel loss function, Stochastic Variance Loss, which explicitly supervises the learning of uncertainty. Results demonstrate consistent performance improvements of about 10\% on the VoxCeleb1-O set and 11\% on the NIST SRE 2024 evaluation set.

说话人识别不确定性建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。