arXiv:2509.16481eess.AScs.SD2025-09中稿 · SPL被引 7

用麦克风间相关性直接分离语音,效果好且计算省。

TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation

  • 用麦克风间的相关性加相位变换做语音分离输入
  • 在LibriCSS数据集上表现优异,计算开销小
  • 适合需要低延迟的实时语音分离场景

多通道语音分离通常利用时频域中麦克风间的相位差(IPDs)与幅度信息拼接,或沿通道轴堆叠实部和虚部。然而,声源的空间信息本质上存在于麦克风之间的差异,特别是它们的相关性;同时每个麦克风的功率也提供了关于源谱的重要信息,因此幅度信息也被保留。为此,我们提出一种直接利用相关性输入并结合相位变换(PHAT)-β来估计分离滤波器的网络。此外,所提出的TF-CorrNet采用双路径策略,交替处理时间与频率轴上的空间信息。进一步地,引入一个频谱模块以建模与源相关的直接时频模式,提升分离性能。实验结果表明,所提TF-CorrNet在LibriCSS数据集上能有效分离语音,在保持高性能的同时具有较低的计算成本。

原文摘要 · Abstract (English)

In general, multi-channel source separation has utilized inter-microphone phase differences (IPDs) concatenated with magnitude information in time-frequency domain, or real and imaginary components stacked along the channel axis. However, the spatial information of a sound source is fundamentally contained in the differences between microphones, specifically in the correlation between them, while the power of each microphone also provides valuable information about the source spectrum, which is why the magnitude is also included. Therefore, we propose a network that directly leverages a correlation input with phase transform (PHAT)-beta to estimate the separation filter. In addition, the proposed TF-CorrNet processes the features alternately across time and frequency axes as a dual-path strategy in terms of spatial information. Furthermore, we add a spectral module to model source-related direct time-frequency patterns for improved speech separation. Experimental results demonstrate that the proposed TF-CorrNet effectively separates the speech sounds, showing high performance with a low computational cost in the LibriCSS dataset.

语音分离相关性建模双路径网络低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。