arXiv:2601.19194eess.AScs.LG2026-01中稿 · ICASSP 2026被引 4

通过自注册语音片段提升多说话人语音识别准确率

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

  • 用说话人活跃时段作为固定条件,增强模型区分力
  • 在EMMA基准上相对原版降低52.4%的平均通话错误率
  • 适合需要跨领域通用的多说话人语音识别场景

多说话人环境下的说话人标注自动语音识别(ASR)仍是重大挑战。尽管某些方法在特定领域微调后表现优异,但跨域泛化能力普遍不足。我们此前提出的分段条件式Whisper(DiCoW)利用说话人分段输出作为条件信息,仅需少量微调即可实现强大的多语言、多领域性能。本文针对DiCoW的关键局限——静音-目标-非目标-重叠(STNO)掩码中的歧义问题(即多个重叠说话人条件相似但转录不同)提出改进。我们引入SE-DiCoW(自注册分段条件式Whisper),利用分段输出定位对话中目标说话人最活跃的录音片段作为固定条件,通过交叉注意力机制在每个编码层注入该条件。同时优化数据分割、模型初始化与增强策略。上述改进使SE-DiCoW在EMMA MT-ASR基准上相较原始DiCoW将宏平均通话错误率(tcpWER)降低了52.4%。

原文摘要 · Abstract (English)

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whisper (DiCoW), leverages speaker diarization outputs as conditioning information and, with minimal fine-tuning, demonstrated strong multilingual and multi-domain performance. In this paper, we address a key limitation of DiCoW: ambiguity in Silence-Target-Non-target-Overlap (STNO) masks, where two or more fully overlapping speakers may have nearly identical conditioning despite differing transcriptions. We introduce SE-DiCoW (Self-Enrolled Diarization-Conditioned Whisper), which uses diarization output to locate an enrollment segment anywhere in the conversation where the target speaker is most active. This enrollment segment is used as fixed conditioning via cross-attention at each encoder layer. We further refine DiCoW with improved data segmentation, model initialization, and augmentation. Together, these advances yield substantial gains: SE-DiCoW reduces macro-averaged tcpWER by 52.4% relative to the original DiCoW on the EMMA MT-ASR benchmark.

语音识别说话人分离Whisper多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。