提出并行估计幅度与相位谱的语音分离模型,提升分离精度。
Neural Speech Separation with Parallel Amplitude and Phase Spectrum Estimation
- 并行建模幅度和相位谱,显式捕捉完整频谱信息
- 在多个数据集上超越时域与隐式相位方法,性能稳定
- 适合需要高保真语音分离的应用场景
本文提出一种新型神经语音分离模型 APSS,采用并行幅度与相位谱估计。不同于多数现有方法,APSS 显式估计相位谱以实现更完整、准确的分离。具体而言,首先从混合语音信号中提取幅度和相位谱,随后通过特征融合器生成联合表示,并由带有时间-频率 Transformer 的深层处理器捕捉时序与谱依赖关系。最后,利用并行的幅度与相位分离器从特征中估计各说话人的对应谱,再通过逆短时傅里叶变换(iSTFT)重建分离语音。实验表明,APSS 在多个数据集上均优于时域分离方法及基于隐式相位估计的时频方法,展现出优异的泛化能力与实际应用潜力。
原文摘要 · Abstract (English)
This paper proposes APSS, a novel neural speech separation model with parallel amplitude and phase spectrum estimation. Unlike most existing speech separation methods, the APSS distinguishes itself by explicitly estimating the phase spectrum for more complete and accurate separation. Specifically, APSS first extracts the amplitude and phase spectra from the mixed speech signal. Subsequently, the extracted amplitude and phase spectra are fused by a feature combiner into joint representations, which are then further processed by a deep processor with time-frequency Transformers to capture temporal and spectral dependencies. Finally, leveraging parallel amplitude and phase separators, the APSS estimates the respective spectra for each speaker from the resulting features, which are then combined via inverse short-time Fourier transform (iSTFT) to reconstruct the separated speech signals. Experimental results indicate that APSS surpasses both time-domain separation methods and implicit-phase-estimation-based time-frequency approaches. Also, APSS achieves stable and competitive results on multiple datasets, highlighting its strong generalization capability and practical applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。