用扩散模型分离语音与背景噪声,无需标注数据却超越有监督方法。
SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
- 基于扩散先验的逆采样框架,联合建模语音与噪声。
- 单麦克风下对1~3人混合语音+噪声均实现更低的字错误率。
- 生成高保真噪声,适合声学场景检测,适用于实际环境应用。
本文针对真实环境下单麦克风语音分离与增强的挑战,提出基于生成逆采样的方法。通过为干净语音和环境噪声分别建立扩散先验,并联合利用它们恢复所有原始信号。我们重构了一种近期的逆采样器以适配该设定。在包含1、2、3名说话人及噪声的混合音频上进行评估,结果表明,尽管完全无监督,本方法在所有条件下均持续优于领先的有监督基线,在字错误率(WER)上表现更优。进一步扩展框架以处理非屏幕内说话人分离。此外,分离出的噪声成分具有高保真度,适用于下游声学场景检测。代码与预训练模型将在论文录用后公开。演示页面:https://ssnaps2026.github.io/ssnaps2026/
原文摘要 · Abstract (English)
This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and ambient noise with dedicated diffusion priors and jointly leverage them to recover all underlying sources. To achieve this, reformulate a recent inverse sampler to match our setting. We evaluate on mixtures of 1, 2, and 3 speakers with noise and show that, despite being entirely unsupervised, our method consistently outperforms leading supervised baselines in WER across all conditions. We further extend our framework to handle off-screen speaker separation. Moreover, the high fidelity of the separated noise component makes it suitable for downstream detection of the acoustic scene. Code and pretrained models will become available upon acceptance. Demo page: https://ssnaps2026.github.io/ssnaps2026/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。