用扩散模型生成式优化语音增强结果,无需重训练。
ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
- 基于噪声空间协方差矩阵,用扩散模型对判别式输出进行生成式修正。
- 在多种环境噪声下显著提升波形与STFT域先进模型的性能。
- 无需训练、兼容多阵列,适合已有语音增强系统快速升级。
多通道语音增强旨在从含噪多通道录音中恢复清晰语音。现有深度学习方法多采用判别式训练,易因回归目标引入非线性失真,尤其在复杂噪声环境下。受无监督多通道源分离方法ArrayDPS启发,本文提出ArrayDPS-Refine,一种利用干净语音扩散先验来优化判别式模型输出的方法。该方法无需训练、生成式且阵列无关:首先从判别式模型的增强语音中估计噪声空间协方差矩阵(SCM),再以该估计的噪声SCM指导扩散后验采样,直接对任意判别式模型输出进行精细化修正。实验表明,ArrayDPS-Refine能持续提升多种判别式模型性能,涵盖当前最先进的波形与频谱域模型。音频演示见 https://xzwy.github.io/ArrayDPSRefineDemo/
原文摘要 · Abstract (English)
Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives, especially under challenging environmental noise conditions. Inspired by ArrayDPS for unsupervised multi-channel source separation, we introduce ArrayDPS-Refine, a method designed to enhance the outputs of discriminative models using a clean speech diffusion prior. ArrayDPS-Refine is training-free, generative, and array-agnostic. It first estimates the noise spatial covariance matrix (SCM) from the enhanced speech produced by a discriminative model, then uses this estimated noise SCM for diffusion posterior sampling. This approach allows direct refinement of any discriminative model's output without retraining. Our results show that ArrayDPS-Refine consistently improves the performance of various discriminative models, including state-of-the-art waveform and STFT domain models. Audio demos are provided at https://xzwy.github.io/ArrayDPSRefineDemo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。