arXiv:2601.09931cs.SD2026-01被引 3

提出新框架,让语音增强同时建模语音和噪声,效果更优且更鲁棒。

Diffusion-based Frameworks for Unsupervised Speech Enhancement

  • 将语音与噪声都作为潜在变量,联合采样提升重建质量。
  • 在WSJ0-QUT和VoiceBank-DEMAND数据集上,性能优于已有无监督方法。
  • 适合追求高音质、低失真语音增强的开发者和研究人员。

本文研究无监督单通道语音增强的基于扩散的方法。先前工作将基于清洁语音训练的分数模型与由非负矩阵分解(NMF)结构化的高斯噪声模型结合,嵌入迭代期望最大化(EM)框架中,通过扩散后验采样估计清洁语音。本文重新审视该框架,提出显式建模语音与声学噪声为潜在变量,在E步中联合采样二者,而非仅采样语音。进一步引入新的半监督框架,以基于扩散的噪声模型替代NMF噪声先验,与语音先验共同学习于单一条件分数模型中。在此框架下,提出两种变体:一种隐式考虑噪声,另一种显式将噪声视为潜在变量。在WSJ0-QUT和VoiceBank-DEMAND数据集上的实验表明,显式噪声建模显著提升性能,无论使用NMF或扩散噪声先验均有效。在匹配条件下,扩散噪声模型在无监督方法中达到最佳整体质量与可懂度;在不匹配条件下,提出的基于NMF的显式噪声框架更具鲁棒性,劣化程度低于多个有监督基线。代码、演示与补充材料已公开。

原文摘要 · Abstract (English)

This paper addresses unsupervised diffusion-based single-channel speech enhancement (SE). Prior work in this direction combines a score-based diffusion model trained on clean speech with a Gaussian noise model whose covariance is structured by non-negative matrix factorization (NMF). This combination is used within an iterative expectation-maximization (EM) scheme, in which a diffusion-based posterior-sampling E-step estimates the clean speech. We first revisit this framework and propose to explicitly model both speech and acoustic noise as latent variables, jointly sampling them in the E-step instead of sampling speech alone as in previous approaches. We then introduce a new semi-supervised SE framework that replaces the NMF noise prior with a diffusion-based noise model, learned jointly with the speech prior in a single conditional score model. Within this framework, we derive two variants: one that implicitly accounts for noise and one that explicitly treats noise as a latent variable. Experiments on WSJ0-QUT and VoiceBank-DEMAND show that explicit noise modeling systematically improves SE performance for both NMF-based and diffusion-based noise priors. Under matched conditions, the diffusion-based noise model attains the best overall quality and intelligibility among unsupervised methods, while under mismatched conditions the proposed NMF-based explicit-noise framework is more robust and suffers less degradation than several supervised baselines. Code, demo, and supplementary materials are publicly available.

语音增强扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。