将说话人特征融入语音识别,提升噪声环境下的识别鲁棒性。
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
- 联合训练语音识别与说话人识别,利用Whisper和ECAPA-TDNN提取多模态特征。
- 在8人嘈杂背景噪声下,性能优于Whisper;对正弦波、噪声语音等增强语音表现更优。
- 适合关注语音识别鲁棒性、多任务模型设计的研究者与工程师。
当前先进的语音识别模型通常将声学信号映射为亚词单元。尽管性能优异,但在背景噪声或语音增强等分布外条件下仍易退化。本文提出,在语音识别中引入说话人表征可提升模型鲁棒性。我们构建了一个基于Transformer的联合模型,同时执行语音识别与说话人识别任务。该模型融合Whisper的语音嵌入与ECAPA-TDNN的说话人嵌入,实现双任务协同。实验表明,该联合模型在干净环境下表现与Whisper相当;在高噪声环境(如8人同时说话的背景噪声)中显著优于Whisper;同时对正弦波语音和噪声合成语音等极端增强语音也表现出更强适应能力。结果表明,将说话人表征整合至语音识别流程,有助于构建更具鲁棒性的模型。
原文摘要 · Abstract (English)
Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background noise and speech augmentations. In this work, we hypothesize that incorporating speaker representations during speech recognition can enhance model robustness to noise. We developed a transformer-based model that jointly performs speech recognition and speaker identification. Our model utilizes speech embeddings from Whisper and speaker embeddings from ECAPA-TDNN, which are processed jointly to perform both tasks. We show that the joint model performs comparably to Whisper under clean conditions. Notably, the joint model outperforms Whisper in high-noise environments, such as with 8-speaker babble background noise. Furthermore, our joint model excels in handling highly augmented speech, including sine-wave and noise-vocoded speech. Overall, these results suggest that integrating voice representations with speech recognition can lead to more robust models under adversarial conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。