arXiv:2501.15965eess.AScs.SD2025-01被引 10

提出一种高效扩散模型,实现单通道语音分离的快速高精度解混。

EDSep: An Effective Diffusion-Based Method for Speech Source Separation

  • 基于随机微分方程设计新型得分匹配方法,提升生成建模效率。
  • 在WSJ0-2mix等数据集上分离效果优于现有扩散与判别模型。
  • 适合追求高保真语音分离性能的研究者或工业应用落地场景。

生成模型在语音分离任务中备受关注,其中基于扩散的方法正被深入探索。尽管扩散技术在生成任务中表现优异,但其在语音分离中的应用仍面临收敛慢、分离效果不佳等挑战。为解决这些问题并提升扩散模型在语音分离中的有效性,本文提出EDSep,一种基于随机微分方程(SDE)得分匹配的单通道语音分离新方法。该方法通过设计新型去噪器函数以更准确逼近数据分布,获得理想去噪输出;同时,精心设计随机采样器,在采样过程中有效求解反向SDE,逐步实现语音源分离。在WSJ0-2mix、LRS2-2mix和VoxCeleb2-2mix等多个数据库上的大量实验表明,EDSep在分离性能上显著优于现有扩散模型及判别式模型,验证了其有效性。

原文摘要 · Abstract (English)

Generative models have attracted considerable attention for speech separation tasks, and among these, diffusion-based methods are being explored. Despite the notable success of diffusion techniques in generation tasks, their adaptation to speech separation has encountered challenges, notably slow convergence and suboptimal separation outcomes. To address these issues and enhance the efficacy of diffusion-based speech separation, we introduce EDSep, a novel single-channel method grounded in score matching via stochastic differential equation (SDE). This method enhances generative modeling for speech source separation by optimizing training and sampling efficiency. Specifically, a novel denoiser function is proposed to approximate data distributions, which obtains ideal denoiser outputs. Additionally, a stochastic sampler is carefully designed to resolve the reverse SDE during the sampling process, gradually separating speech from mixtures. Extensive experiments on databases such as WSJ0-2mix, LRS2-2mix, and VoxCeleb2-2mix demonstrate our proposed method's superior performance over existing diffusion and discriminative models, validating its efficacy.

语音分离扩散模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。