arXiv:2603.24810eess.AS2026-03

用扩散模型统一提升多麦克风语音增强与分离效果

Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation

  • 以扩散模型为先验,对已有模型输出进行生成式优化
  • 无需训练,可适配不同阵列、模型和任务,显著降低失真
  • 适合需要自然语音输出的语音系统开发者

我们提出Uni-ArrayDPS,一种基于扩散模型的统一多通道语音增强与分离精炼框架。现有方法多为判别式,虽能生成高信噪比输出,但仍可能引入非线性失真。为此,我们提出使用语音扩散先验,对任意强判别模型的输出进行精炼。该方法无需训练,具备阵列无关性与任务通用性,支持增强与分离。给定判别模型输出与噪声混合信号,我们估计噪声空间协方差矩阵(SCM),用于计算扩散后验采样所需的似然。仅需预训练的干净语音扩散模型作为先验,不需额外训练或微调,即可直接跨任务、跨阵列几何与模型架构泛化。大量实验表明,Uni-ArrayDPS在多种判别模型上持续提升增强与分离性能,并在真实数据集上取得优异结果。音频示例见 https://xzwy.github.io/Uni-ArrayDPS/

原文摘要 · Abstract (English)

We propose Uni-ArrayDPS, a novel diffusion-based refinement framework for unified multi-channel speech enhancement and separation. Existing methods for multi-channel speech enhancement/separation are mostly discriminative and are highly effective at producing high-SNR outputs. However, they can still generate unnatural speech with non-linear distortions caused by the neural network and regression-based objectives. To address this issue, we propose Uni-ArrayDPS, which refines the outputs of any strong discriminative model using a speech diffusion prior. Uni-ArrayDPS is generative, array-agnostic, and training-free, and supports both enhancement and separation. Given a discriminative model's enhanced/separated speech, we use it, together with the noisy mixtures, to estimate the noise spatial covariance matrix (SCM). We then use this SCM to compute the likelihood required for diffusion posterior sampling of the clean speech source(s). Uni-ArrayDPS requires only a pre-trained clean-speech diffusion model as a prior and does not require additional training or fine-tuning, allowing it to generalize directly across tasks (enhancement/separation), microphone array geometries, and discriminative model backbones. Extensive experiments show that Uni-ArrayDPS consistently improves a wide range of discriminative models for both enhancement and separation tasks. We also report strong results on a real-world dataset. Audio demos are provided at \href{https://xzwy.github.io/Uni-ArrayDPS/}{https://xzwy.github.io/Uni-ArrayDPS/}.

语音增强扩散模型多通道生成式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。