arXiv:2501.15466eess.AScs.SD2025-01被引 2

新模型让语音识别在嘈杂和重叠录音下仍保持高精度。

End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario

  • 用双注意力机制从混乱录音中提取说话人特征。
  • 在5dB信干比下重叠录音时错误率仅16.44%,远超传统方法。
  • 适合实际场景中不完美录音的语音助手应用。

本文提出一种新型流式端到端目标说话人语音识别模型,解决现有系统在噪声录入和特定录音要求下的两大瓶颈。该模型采用带双重注意力机制的靶向说话人循环神经网络转换器(TS-RNNT),实现上下文引导与重叠录入处理。文本解码器与注意力机制专门设计用于从含噪、重叠的录入音频中提取有效说话人特征。在合成数据集上的实验表明,该模型具备强鲁棒性,在5dB信干比(SIR)的重叠录入条件下,词错误率(WER)维持在16.44%,而传统方法在此条件下错误率超过75%。这一显著性能提升,结合半文本依赖的录入能力,标志着语音控制设备向更实用、更灵活方向的重要进展。

原文摘要 · Abstract (English)

This paper presents a novel streaming end-to-end target-speaker speech recognition that addresses two critical limitations in systems: the handling of noisy enrollment utterances and specific enrollment phrase requirements. This paper proposes a robust Target-Speaker Recurrent Neural Network Transducer (TS-RNNT) with dual attention mechanisms for contextual biasing and overlapping enrollment processing. The model incorporates a text decoder and attention mechanism specifically designed to extract relevant speaker characteristics from noisy, overlapping enrollment audio. Experimental results on a synthesized dataset demonstrate the model's resilience, maintaining a Word Error Rate (WER) of 16.44% even with overlapping enrollment at 5dB Signal-to-Interference Ratio (SIR), compared to conventional approaches that degrade to WERs above 75% under similar conditions. This significant performance improvement, coupled with the model's semi-text-dependent enrollment capabilities, represents a substantial advancement toward more practical and versatile voice-controlled devices.

语音识别说话人识别端到端降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。