arXiv:2507.09226eess.AScs.SD2025-07中稿 · IEEE ASRU 2025被引 1

用语音识别数据训练说话人分离模型效果差,紧对齐能显著提升性能。

Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?

  • 用强制对齐生成严格边界,替代松散标注的ASR数据
  • 紧边界训练使流式场景下的错误率降低23.4%(相对减少)
  • 提升说话人分离性能,同时轻微改善语音识别结果

神经网络说话人分离广泛用于重叠语音场景,但需要大量多说话人数据训练。当前常通过合并多个语料库构建数据集,包括原本为多说话人语音识别(ASR)设计的数据。然而,这些ASR数据通常存在边界模糊的问题,与说话人分离基准要求的严格边界不一致。本文发现,这种边界松散会显著增加分离误差,降低评估可靠性。模型在不同边界精度数据上训练后,会学习到特定数据集的松散模式,导致跨数据集泛化能力差。通过强制对齐生成标准化紧边界进行训练,不仅提升了说话人分离性能(尤其在流式场景中),还通过简单后处理略微改善了语音识别表现。

原文摘要 · Abstract (English)

Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora, including those originally designed for multi-speaker automatic speech recognition (ASR). However, ASR datasets often feature loosely defined segment boundaries that do not align with the stricter conventions of diarization benchmarks. In this work, we show that such boundary looseness significantly impacts the diarization error rate, reducing evaluation reliability. We also reveal that models trained on data with varying boundary precision tend to learn dataset-specific looseness, leading to poor generalization across out-of-domain datasets. Training with standardized tight boundaries via forced alignment improves not only diarization performance, especially in streaming scenarios, but also ASR performance when combined with simple post-processing.

说话人分离语音识别数据对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。