分离与识别解耦,让多说话人语音识别更准。
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
- 分离和识别分开训练,识别器只用干净语音
- 在Libri2Mix上达到5.1%的词错误率
- 适合需要高精度多说话人识别的场景
尽管深度学习推动了自动语音识别(ASR)的巨大进展,但在真实世界的多说话人场景中表现仍不理想。说话人分离虽能有效分离各说话人,但作为前端会引入处理伪影,损害在纯净语音上训练的ASR后端性能。主流方法通过在噪声语音上训练后端以规避伪影问题。本文提出将说话人分离前端与ASR后端的训练解耦,后端仅在干净语音上训练。该解耦系统在Libri2Mix开发集/测试集上取得5.1%的词错误率(WER),显著优于其他多说话人ASR基线。在1通道和6通道SMS-WSJ上分别达到7.60%/5.74%的先进水平。在录制的LibriCSS数据集上,实现2.92%的说话人归属词错误率。这些最先进结果表明,解耦分离与识别是提升鲁棒多说话人ASR的有效策略。
原文摘要 · Abstract (English)
Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing artifacts that degrade the ASR backend trained on clean speech. As a result, mainstream robust ASR systems train the backend on noisy speech to avoid processing artifacts. In this work, we propose to decouple the training of the speaker separation frontend and the ASR backend, with the latter trained on clean speech only. Our decoupled system achieves 5.1% word error rates (WER) on the Libri2Mix dev/test sets, significantly outperforming other multi-talker ASR baselines. Its effectiveness is also demonstrated with the state-of-the-art 7.60%/5.74% WERs on 1-ch and 6-ch SMS-WSJ. Furthermore, on recorded LibriCSS, we achieve the speaker-attributed WER of 2.92%. These state-of-the-art results suggest that decoupling speaker separation and recognition is an effective approach to elevate robust multi-talker ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。