arXiv:2606.02327eess.AS2026-06

通过估计人工噪声,让降噪模型在真实噪声下也能有效工作。

Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets

  • 同时估计自然噪声和人工噪声,用差值消除残留噪声
  • 在WHAM!和CHiME-3数据集上显著提升去噪效果
  • 兼容传统训练数据,适合实际场景的弱监督应用

语音降噪是人耳听觉和下游系统处理嘈杂现实声学环境的重要步骤。然而,由于无法人工标注干净语音作为训练目标(因为生成干净版本本身就是任务),传统的域内监督训练难以实现。通常采用在干净语音中添加人工噪声的方法,但这些噪声来自受控领域,导致神经网络泛化能力差。另一种方法是使用真实噪声录音作为目标(即噪点目标训练,NyTT),但其优化目标并非纯净语音。本文发现:若同时估计自然噪声与人工噪声,可利用残余噪声的不分离性——通过简单相减即可抵消残留噪声。关键在于该最优解与传统人工混合数据一致,支持两类数据联合训练,提升模型对真实场景的适应能力。实验在WHAM!和CHiME-3基准上验证了该方法的有效性。

原文摘要 · Abstract (English)

Speech denoising is an often necessary step not only for human listening, but also for downstream processing by systems lacking robustness to noisy, real-world acoustic conditions. Unfortunately, denoising is a problem where conventional in-domain supervised training is not trivial, as the training targets cannot be annotated by humans: producing a clean version of a naturally-noisy speech recording is itself the task to solve. Supervised training is typically performed through the artificial addition of noise to clean speech recordings, which can only be sourced from controlled domains, a significant limitation due to the poor out-of-domain generalization of neural networks. An alternative is noisy target training (NyTT), which simply replaces the clean speech with in-domain noisy recordings, with the hope that learning to remove the artificial noise will extend to the natural. Though having shown promising results, NyTT's training objective is not minimized by clean speech estimates. We show that by estimating the artificial noise in addition to the naturally-noisy speech, the undesirable optimum can actually be exploited: the residual noise in the speech estimate can be canceled by the noise estimate via simple subtraction. Crucially, the optimum is fully compatible with conventional artificial mixtures, enabling joint training using both types of data with consistent optimization targets, opening the door to improved domain adaptability. The effectiveness of our approach is demonstrated through WHAM! and CHiME-3-based benchmarks.

语音降噪弱监督噪声建模域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。