arXiv:2506.12672cs.SDcs.CL2025-06中稿 · Interspeech 2025

让语音识别模型知道谁在何时说话,提升多人重叠对话的识别准确率。

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

  • 用说话人嵌入和活动信息显式指导解码器聚焦目标说话人
  • 在多说话人重叠场景下显著提升识别效果,无需额外录音注册
  • 结合端到端语音分离模型,实现无须声纹注册的实时说话人区分

我们提出一种基于序列输出训练的增强方法——说话人条件化序列输出训练(SC-SOT),用于端到端多说话人语音识别。首先分析SOT对重叠语音的处理机制,发现解码器具备隐式的说话人分离能力。但因重叠区域声学线索模糊,该能力常不足。为此,SC-SOT显式地将说话人信息引入解码器,提供“谁在何时说话”的详细信息。具体通过:(1) 说话人嵌入,使模型聚焦目标说话人的声学特征;(2) 说话人活跃度信息,引导模型抑制非目标说话人。说话人嵌入来自联合训练的端到端说话人分组模型,避免了传统声纹注册需求。实验表明,该条件化方法在重叠语音场景中显著有效。

原文摘要 · Abstract (English)

We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and we found the decoder performs implicit speaker separation. We hypothesize this implicit separation is often insufficient due to ambiguous acoustic cues in overlapping regions. To address this, SC-SOT explicitly conditions the decoder on speaker information, providing detailed information about "who spoke when". Specifically, we enhance the decoder by incorporating: (1) speaker embeddings, which allow the model to focus on the acoustic characteristics of the target speaker, and (2) speaker activity information, which guides the model to suppress non-target speakers. The speaker embeddings are derived from a jointly trained E2E speaker diarization model, mitigating the need for speaker enrollment. Experimental results demonstrate the effectiveness of our conditioning approach on overlapped speech.

语音识别多说话人端到端说话人分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。