arXiv:2508.18833eess.AS2025-08中稿 · 16th ITG Conferenc…

用扩散模型同时去噪回声,效果优于单一处理。

On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation

  • 分步使用去噪与去混响模型,按主导失真顺序效果最佳。
  • 联合训练三类数据的单模型,在多场景下表现最均衡。
  • 适合实际语音增强应用,尤其在噪声与混响并存时。

扩散模型在降噪或去混响语音增强中已展现自然音质效果。然而,其同时处理噪声与混响的能力尚未被充分研究,而这种情况在实际应用中最常见。本文探讨了多种增强噪声和/或混响语音的方法:一是分别训练去噪与去混响模型后级联使用;二是训练单一模型,该模型仅针对纯噪声、纯混响及噪声混响混合数据。实验基于人工生成和真实录音数据进行。结果表明,级联模型需按主导失真顺序应用才能达到满意效果;若需一个能处理所有情况的统一模型,则在三类数据(纯噪声、纯混响、噪声混响)上联合训练的模型表现最优。

原文摘要 · Abstract (English)

Diffusion models have been shown to achieve natural-sounding enhancement of speech degraded by noise or reverberation. However, their simultaneous denoising and dereverberation capability has so far not been studied much, although this is arguably the most common scenario in a practical application. In this work, we investigate different approaches to enhance noisy and/or reverberant speech. We examine the cascaded application of models, each trained on only one of the distortions, and compare it with a single model, trained either solely on data that is both noisy and reverberated, or trained on data comprising subsets of purely noisy, of purely reverberated, and of noisy reverberant speech. Tests are performed both on artificially generated and real recordings of noisy and/or reverberant data. The results show that, when using the cascade of models, satisfactory results are only achieved if they are applied in the order of the dominating distortion. If only a single model is desired that can operate on all distortion scenarios, the best compromise appears to be a model trained on the aforementioned three subsets of degraded speech data.

扩散模型语音增强去混响去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。