arXiv:2506.12500eess.AScs.SD2025-06中稿 · Interspeech 2025

改进语音嵌入模型在多说话人场景下的偏置问题,提升识别准确率。

Mitigating Non-Target Speaker Bias in Guided Speaker Embedding

  • 利用目标说话人活跃区间计算统计量,减少非目标干扰。
  • 在高低重叠场景下均提升说话人验证与聚类性能。
  • 适用于真实对话中重叠频繁的语音识别任务。

在多说话人场景中获取高质量说话人嵌入对诸多应用至关重要。近期提出的引导式说话人嵌入框架通过利用目标与非目标说话人的语音活动作为线索,在严重重叠情况下显著提升了嵌入质量,仅在低重叠情况下有小幅退化。然而,由于自然对话中极端重叠情况罕见,这种退化不容忽视。本文首次揭示,该退化源于广泛使用的基于全局统计的模块对仅含非目标说话人的时间段过于敏感。为此,我们提出扩展此类模块,利用目标说话人活跃线索,从目标活跃区间计算统计量。所提方法在低与高重叠比下均提升了说话人验证性能,并在多个数据集上改善了说话人分离性能。

原文摘要 · Abstract (English)

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets.

说话人嵌入语音识别去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。