用语音分角色模型识别并清除多说话人语音识别中的混淆片段,提升准确率。
Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction

- 基于分角色模型的三重验证机制,识别并移除错误归属的语音片段。
- 在高重叠场景下,相对减少29%的字符级词错误率(cpW ER)。
- 适合需要高精度多说话人语音转录的会议记录、法庭记录等场景。
尽管级联式多说话人语音识别(MT-ASR)利用了先进的基础模型,其性能常受限于分离过程中的说话人泄漏问题。以往的修正策略主要聚焦于词汇层面的说话人重新标注。本文提出一种互补的剪枝范式,可鲁棒地识别并去除泄漏伪影。该方法利用预训练的说话人分角色模型作为多模态验证器,通过时间包含、词汇交叉验证与时间对齐的三重共识来剪枝转录段。在LibriMix、LibriSpeechMix和AMI Meeting数据集上的实验表明,该算法在多种重叠条件下均能持续降低字符级词错误率(cpW ER)。特别是在高说话人泄漏子集上,相对减少了最高达29%的cpW ER,凸显其在复杂声学环境中提升级联式MT-ASR转录可靠性的有效性。
原文摘要 · Abstract (English)
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identifies and removes leakage artifacts. Our method utilizes a pre-trained speaker diarization model as a multimodal verifier to prune transcribed segments satisfying a tripartite consensus of temporal containment, lexical cross-validation, and temporal alignment. Results on LibriMix, LibriSpeechMix, and the AMI Meeting corpus show our algorithm consistently reduces cpW ER across diverse overlap conditions. Specifically, on subsets with high speaker leakage, our method achieves relative cpW ER reductions of up to 29%, highlighting its effectiveness in enhancing the reliability of cascaded MT-ASR transcripts in complex acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。