融合语音分离与识别模型,提升多说话人语音识别准确率
BUT System for the MLC-SLM Challenge
- 用改进的DiCoW和DiariZen系统实现说话人分离与识别
- 微平均tcpWER/CER达16.75%,在挑战赛中排名第二
- 发现数据标注问题并提出简化修复策略,增强模型鲁棒性
我们提出一个双说话人自动语音识别系统,结合DiCoW(基于Whisper的说话人条件化模型)与DiariZen(基于Pyannote的语音分离流水线)。在无领域外(OOD)多语言场景下,未微调的DiariZen持续优于基准Pyannote模型,表现出强泛化能力。尽管DiCoW仅在英语数据上微调用于目标说话人识别,仍保持良好多语言性能,说明编码器修改未损害Whisper的多语言能力。随后,在MLC-SLM挑战数据上对两者进行微调,细调后的DiariZen继续优于细调版Pyannote基线,而DiCoW则通过领域适应获得进一步提升。最终系统在任务2中取得微平均tcpWER/CER 16.75%,排名第二。最后,我们发现训练数据存在标注不一致问题,如缺失语音段和错误静音标注,提出简单缓解策略以改善模型鲁棒性。
原文摘要 · Abstract (English)
We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline built on top of Pyannote. We first evaluate both systems in out-of-domain (OOD) multilingual scenarios without any fine-tuning. In this scenario, DiariZen consistently outperforms the baseline Pyannote diarization model, demonstrating strong generalization. Despite being fine-tuned on English-only data for target-speaker ASR, DiCoW retains solid multilingual performance, indicating that encoder modifications preserve Whisper's multilingual capabilities. We then fine-tune both DiCoW and DiariZen on the MLC-SLM challenge data. The fine-tuned DiariZen continues to outperform the fine-tuned Pyannote baseline, while DiCoW sees further gains from domain adaptation. Our final system achieves a micro-average tcpWER/CER of 16.75% and ranks second in Task 2 of the MLC-SLM challenge. Lastly, we identify several labeling inconsistencies in the training data -- such as missing speech segments and incorrect silence annotations -- which can hinder diarization fine-tuning. We propose simple mitigation strategies to address these issues and improve system robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。