利用目标说话人语音特征提升嘈杂环境下的方向定位精度。
Robust Target Speaker Direction of Arrival Estimation
- 用目标说话人已录语音作参考,融合全带与子带频谱信息。
- 在LibriSpeech数据集上实现更优的多说话人场景定位性能。
- 适合语音增强、智能麦克风等需要精准声源定位的应用。
在多说话人环境中,目标说话人的方向到达(DOA)对提升语音清晰度和提取目标语音至关重要。然而,传统DOA估计方法在噪声、混响及竞争说话人存在时表现不佳。为此,我们提出RTS-DOA系统,一种鲁棒且实时的DOA估计方案。该系统创新性地使用目标说话人注册语音作为参考,并结合麦克风阵列的全带与子带频谱信息进行方向估计。系统包含语音增强模块以初步提升语音质量,空间模块用于学习空间特征,以及说话人模块提取语音指纹特征。在LibriSpeech数据集上的实验表明,RTS-DOA能有效应对多说话人场景,建立了新的性能基准。
原文摘要 · Abstract (English)
In multi-speaker environments the direction of arrival (DOA) of a target speaker is key for improving speech clarity and extracting target speaker's voice. However, traditional DOA estimation methods often struggle in the presence of noise, reverberation, and particularly when competing speakers are present. To address these challenges, we propose RTS-DOA, a robust real-time DOA estimation system. This system innovatively uses the registered speech of the target speaker as a reference and leverages full-band and sub-band spectral information from a microphone array to estimate the DOA of the target speaker's voice. Specifically, the system comprises a speech enhancement module for initially improving speech quality, a spatial module for learning spatial information, and a speaker module for extracting voiceprint features. Experimental results on the LibriSpeech dataset demonstrate that our RTS-DOA system effectively tackles multi-speaker scenarios and established new optimal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。