新模型让语音识别在嘈杂和重叠录音下仍保持高精度。
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
- 用双注意力机制从混乱录音中提取说话人特征。
- 在5dB信干比下重叠录音时错误率仅16.44%,远超传统方法。
- 适合实际场景中不完美录音的语音助手应用。
本文提出一种新型流式端到端目标说话人语音识别模型,解决现有系统在噪声录入和特定录音要求下的两大瓶颈。该模型采用带双重注意力机制的靶向说话人循环神经网络转换器(TS-RNNT),实现上下文引导与重叠录入处理。文本解码器与注意力机制专门设计用于从含噪、重叠的录入音频中提取有效说话人特征。在合成数据集上的实验表明,该模型具备强鲁棒性,在5dB信干比(SIR)的重叠录入条件下,词错误率(WER)维持在16.44%,而传统方法在此条件下错误率超过75%。这一显著性能提升,结合半文本依赖的录入能力,标志着语音控制设备向更实用、更灵活方向的重要进展。
原文摘要 · Abstract (English)
This paper presents a novel streaming end-to-end target-speaker speech recognition that addresses two critical limitations in systems: the handling of noisy enrollment utterances and specific enrollment phrase requirements. This paper proposes a robust Target-Speaker Recurrent Neural Network Transducer (TS-RNNT) with dual attention mechanisms for contextual biasing and overlapping enrollment processing. The model incorporates a text decoder and attention mechanism specifically designed to extract relevant speaker characteristics from noisy, overlapping enrollment audio. Experimental results on a synthesized dataset demonstrate the model's resilience, maintaining a Word Error Rate (WER) of 16.44% even with overlapping enrollment at 5dB Signal-to-Interference Ratio (SIR), compared to conventional approaches that degrade to WERs above 75% under similar conditions. This significant performance improvement, coupled with the model's semi-text-dependent enrollment capabilities, represents a substantial advancement toward more practical and versatile voice-controlled devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。