利用方向与距离线索提升麦克风阵列的语音分离效果
Exploring Efficient Directional and Distance Cues for Regional Speech Separation
- 结合改进的延时求和法获取方向信息,增强目标方向信号
- 引入直达比作为特征,显著提升近距与远距语音的分离性能
- 在真实对话场景下表现优异,适合实际应用部署
本文提出一种基于神经网络的麦克风阵列区域语音分离方法,利用新型空间线索同时提取特定方向和指定距离内的声源。该方法改进了延时求和技术以获取更精准的方向性提示,显著增强目标方向信号;进一步将直达比(direct-to-reverberant ratio)融入输入特征,使模型能够更好区分近距与远距声源。实验表明,该方法在多个客观指标上均取得显著提升,并在真实场景录制的CHiME-8 MMCSG数据集上达到当前最优性能,验证了其在实际应用中的有效性。
原文摘要 · Abstract (English)
In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound source not only from specified direction but also within defined distance. Specifically, our method employs an improved delay-and-sum technique to obtain directional cues, substantially enhancing the signal from the target direction. We further enhance separation by incorporating the direct-to-reverberant ratio into the input features, enabling the model to better discriminate sources within and beyond a specified distance. Experimental results demonstrate that our proposed method leads to substantial gains across multiple objective metrics. Furthermore, our method achieves state-of-the-art performance on the CHiME-8 MMCSG dataset, which was recorded in real-world conversational scenarios, underscoring its effectiveness for speech separation in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。