arXiv:2603.29097eess.AScs.SD2026-03被引 2

通过时频相关性建模,实现语音分离的渐进式优化。

Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation

  • 采用非对称编解码结构,分阶段提升语音分离效果。
  • 在多个数据集上实现单/多通道下的稳定性能提升。
  • 适合需要高鲁棒性的实际语音分离场景使用。

真实声学环境中的语音分离仍具挑战性,因重叠说话人、背景噪声和混响需同时处理。尽管近期时频(TF)域模型表现优异,但多数仍依赖晚期分割架构,将说话人解耦推迟至最后阶段,造成信息瓶颈,在恶劣条件下削弱判别能力。为此,我们提出SR-CorrNet,一种基于时频相关性的非对称编码器-解码器框架,将分离-重建(SepRe)策略引入TF双路径主干网络。编码器从混合信号中完成粗分离,共享权重的解码器通过跨说话人间交互逐步重构具有说话人判别性的特征,实现阶段式优化。为配合该架构,我们将语音分离建模为结构化相关性到滤波器的问题:利用观测信号计算出的时空谱相关性作为输入特征,估计深层滤波器以恢复目标信号。进一步引入基于吸引子的动态分割模块,自适应调整输出流数量以匹配实际说话人数。在WSJ0-{2,3,4,5}Mix、WHAMR!和LibriCSS上的实验结果表明,无论在无混响、噪声混响还是真实录音条件下,单通道与多通道设置下均取得持续改进,验证了基于相关性的滤波器估计在时频域中进行分离-重建的有效性。

原文摘要 · Abstract (English)

Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong performance, most still rely on late-split architectures, where speaker disentanglement is deferred to the final stage, creating an information bottleneck and weakening discriminability under adverse conditions. To address this issue, we propose SR-CorrNet, an asymmetric encoder-decoder framework that introduces the separation-reconstruction (SepRe) strategy into a TF dual-path backbone. The encoder performs coarse separation from mixture observations, while the weight-shared decoder progressively reconstructs speaker-discriminative features with cross-speaker interaction, enabling stage-wise refinement. To complement this architecture, we formulate speech separation as a structured correlation-to-filter problem: spatio-spectro-temporal correlations computed from the observations are used as input features, and the corresponding deep filters are estimated to recover target signals. We further incorporate an attractor-based dynamic split module to adapt the number of output streams to the actual speaker configuration. Experimental results on WSJ0-{2,3,4,5}Mix, WHAMR!, and LibriCSS demonstrate consistent improvements across anechoic, noisy-reverberant, and real-recorded conditions in both single- and multi-channel settings, highlighting the effectiveness of TF-domain SepRe with correlation-based filter estimation for speech separation.

语音分离时频分析深度学习非对称结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。