arXiv:2508.06393cs.SDcs.AI2025-08中稿 · Interspeech 2025被引 1

无需提前注册说话人,实现噪声下高精度语音分离与辨识

Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling

  • 通过自动采样目标说话人嵌入,实现无注册训练
  • 在重叠语音段提升识别准确率,DER降低71%,cpWER降69%
  • 适合噪声环境下的多人对话分析与语音系统部署

传统语音分离与说话人辨识方法依赖目标说话人的先验知识或参与者数量的预设。为克服这些局限,近期研究转向无需注册的方法,可在无显式标注情况下识别目标说话人。本文提出一种新方法,利用混合语音中自动识别的目标说话人嵌入,训练同步语音分离与辨识模型。所提模型采用双阶段训练流程,学习对背景噪声具有鲁棒性的说话人表示特征。此外,设计了一种针对重叠语音帧的重叠谱损失函数,专门提升辨识精度。实验结果表明,相比当前最优基线,该方法在 DER 上实现71%相对提升,在 cpWER 上提升69%。

原文摘要 · Abstract (English)

Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus on developing enrollment-free methods capable of identifying targets without explicit speaker labeling. This work introduces a new approach to train simultaneous speech separation and diarization using automatic identification of target speaker embeddings, within mixtures. Our proposed model employs a dual-stage training pipeline designed to learn robust speaker representation features that are resilient to background noise interference. Furthermore, we present an overlapping spectral loss function specifically tailored for enhancing diarization accuracy during overlapped speech frames. Experimental results show significant performance gains compared to the current SOTA baseline, achieving 71% relative improvement in DER and 69% in cpWER.

说话人辨识语音分离鲁棒性重叠语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。