用定向噪声引导的轻量扩散模型,高效提升语音降噪效果
GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
- 引入各向异性噪声引导,保留原始语音中的清晰成分
- 仅需约450万参数,性能超越现有扩散模型
- 适合资源受限设备,在极噪环境下表现优异
语音增强旨在提升复杂噪声环境下语音的可懂度和质量。近年来,扩散模型在该领域受到广泛关注,取得了良好效果。现有基于扩散的方法通常使用各向同性高斯噪声对信号进行模糊处理,并从先验中恢复干净语音,但计算开销大。本文认为,语音增强本质上并非纯生成任务,而更侧重于噪声抑制与缺失信息补全,原混合信号中的干净线索无需重新生成。为此,提出一种在扩散过程中引入各向异性引导噪声的方法,使神经网络能够保留噪声录音中的干净成分。该方法显著降低计算复杂度,同时在多种噪声和语音失真条件下表现出强鲁棒性。实验表明,所提方法仅需约450万参数,即达到当前最优性能,大幅缩小了扩散模型与传统预测式语音增强方法之间的模型规模差距。此外,该方法在极端嘈杂场景中依然表现良好,具备在严苛环境应用的潜力。
原文摘要 · Abstract (English)
Speech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion model has gained lots of attention in speech enhancement area, achieving competitive results. Current diffusion-based methods blur the signal with isotropic Gaussian noise and recover clean speech from the prior. However, these methods often suffer from a substantial computational burden. We argue that the computational inefficiency partially stems from the oversight that speech enhancement is not purely a generative task; it primarily involves noise reduction and completion of missing information, while the clean clues in the original mixture do not need to be regenerated. In this paper, we propose a method that introduces noise with anisotropic guidance during the diffusion process, allowing the neural network to preserve clean clues within noisy recordings. This approach substantially reduces computational complexity while exhibiting robustness against various forms of noise and speech distortion. Experiments demonstrate that the proposed method achieves state-of-the-art results with only approximately 4.5 million parameters, a number significantly lower than that required by other diffusion methods. This effectively narrows the model size disparity between diffusion-based and predictive speech enhancement approaches. Additionally, the proposed method performs well in very noisy scenarios, demonstrating its potential for applications in highly challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。