arXiv:2507.20926eess.AS2025-07中稿 · INTERSPEECH 2025被引 5

利用声音方向信息精准提取指定方向语音,提升嘈杂环境下的语音清晰度。

End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios

  • 通过方向与波束宽度嵌入,端到端定位目标说话人方向
  • 在多说话人混杂场景中显著提升目标语音信噪比
  • 特别适合需要高精度语音识别的现实应用

目标说话人提取(TSE)在噪声和多说话人环境中对增强语音信号至关重要。本文提出一种端到端的TSE模型,融合到达方向(DOA)与波束宽度嵌入,从指定空间区域中提取语音。该方法高效捕捉时空特征,在多个说话人同时发声的复杂场景下表现稳健。实验表明,所提模型不仅能显著增强目标语音,还能有效抑制其他方向干扰,生成清晰孤立的目标语音。此外,该模型在下游自动语音识别(ASR)任务中取得显著性能提升,适用于真实应用场景。

原文摘要 · Abstract (English)

Target Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorporates Direction of Arrival (DOA) and beamwidth embeddings to extract speech from a specified spatial region centered around the DOA. Our approach efficiently captures spatial and temporal features, enabling robust performance in highly complex scenarios with multiple simultaneous speakers. Experimental results demonstrate that the proposed model not only significantly enhances the target speech within the defined beamwidth but also effectively suppresses interference from other directions, producing a clear and isolated target voice. Furthermore, the model achieves remarkable improvements in downstream Automatic Speech Recognition (ASR) tasks, making it particularly suitable for real-world applications.

语音分离方向感知端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。