arXiv:2412.08856cs.SDeess.AS2024-12被引 1

通过复数循环一致性机制,提升单声道语音增强的音质。

Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement

  • 分两阶段估计语音谱的幅度与相位,引入噪声感知逆过程。
  • 利用幅度与相位的内在关联,实现更精准的语音重建。
  • 适合语音增强、降噪领域研究者,尤其关注相位建模的场景。

本文提出一种基于扩散模型的单声道语音增强方法。该方法在两个扩散网络中分别估计语音谱的幅度和相位,并在扩散过程中逐步添加真实世界噪声片段,提出噪声感知的逆过程以学习生成干净语音谱和噪声谱。为充分利用幅度与相位之间的内在关系,引入复数循环一致性(CCC)机制,使估计的幅度映射相位,反之亦然。该算法在相位感知语音增强扩散模型(SEDM)中实现。在公开数据集上的大量实验表明,利用幅度与相位的内在关系能显著提升语音质量,相比传统扩散模型具有明显优势。

原文摘要 · Abstract (English)

In this paper, we present a novel diffusion model-based monaural speech enhancement method. Our approach incorporates the separate estimation of speech spectra's magnitude and phase in two diffusion networks. Throughout the diffusion process, noise clips from real-world noise interferences are added gradually to the clean speech spectra and a noise-aware reverse process is proposed to learn how to generate both clean speech spectra and noise spectra. Furthermore, to fully leverage the intrinsic relationship between magnitude and phase, we introduce a complex-cycle-consistent (CCC) mechanism that uses the estimated magnitude to map the phase, and vice versa. We implement this algorithm within a phase-aware speech enhancement diffusion model (SEDM). We conduct extensive experiments on public datasets to demonstrate the effectiveness of our method, highlighting the significant benefits of exploiting the intrinsic relationship between phase and magnitude information to enhance speech. The comparison to conventional diffusion models demonstrates the superiority of SEDM.

语音增强扩散模型相位建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。