arXiv:2502.06285cs.SDcs.AI2025-02被引 5

利用相对传递函数实现多麦克风语音分离,效果优于传统方法。

End-to-End Multi-Microphone Speaker Extraction Using Relative Transfer Functions

  • 用参考语音估计瞬时相对传递函数作为空间线索
  • 在混响环境中分离目标说话人,性能优于方向到达和频谱嵌入法
  • 适合需要高精度语音增强的智能会议、助听设备场景

本文提出一种多麦克风方法,用于从包含多个说话人和定向噪声的混响环境中提取目标说话人。该方法利用从与目标声源相同位置录制的参考语音中估计出的瞬时相对传递函数(RTF)作为空间线索。在复杂声学场景下的实验结果表明,使用空间线索的效果优于基于频谱的线索,且瞬时RTF的表现优于基于方向到达(DOA)的空间线索。

原文摘要 · Abstract (English)

This paper introduces a multi-microphone method for extracting a desired speaker from a mixture involving multiple speakers and directional noise in a reverberant environment. In this work, we propose leveraging the instantaneous relative transfer function (RTF), estimated from a reference utterance recorded in the same position as the desired source. The effectiveness of the RTF-based spatial cue is compared with direction of arrival (DOA)-based spatial cue and the conventional spectral embedding. Experimental results in challenging acoustic scenarios demonstrate that using spatial cues yields better performance than the spectral-based cue and that the instantaneous RTF outperforms the DOA-based spatial cue.

语音分离多麦克风空间线索混响环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。