arXiv:2410.05920cs.SDcs.AI2024-10NeurIPS被引 28

用GAN生成接近录音棚品质的清晰语音,支持真实场景降噪。

FINALLY: fast and universal speech enhancement with studio-like quality

  • 基于WavLM感知损失与MS-STFT对抗训练,提升模型稳定性。
  • 在多个数据集上实现48kHz高质量语音输出,性能达当前最优。
  • 适合语音增强、音频修复等需要高保真输出的应用场景。

本文针对真实录音中常见的背景噪声、混响和麦克风畸变等问题,重新审视生成对抗网络(GAN)在语音增强中的应用。理论证明GAN天然倾向于寻找条件清洁语音分布中的最大密度点,这正是语音增强任务的核心需求。通过研究多种感知损失特征提取器,我们提出一种探测特征空间结构的方法,并将基于WavLM的感知损失融入MS-STFT对抗训练流程,构建出稳定高效的训练方案。由此提出的模型FINALLY基于HiFi++架构,引入WavLM编码器与全新训练流程。在多个数据集上的实验证明,该模型可在48 kHz下生成清晰高质语音,达到语音增强领域的最先进水平。

原文摘要 · Abstract (English)

In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the use of Generative Adversarial Networks (GANs) for speech enhancement and theoretically show that GANs are naturally inclined to seek the point of maximum density within the conditional clean speech distribution, which, as we argue, is essential for the speech enhancement task. We study various feature extractors for perceptual loss to facilitate the stability of adversarial training, developing a methodology for probing the structure of the feature space. This leads us to integrate WavLM-based perceptual loss into MS-STFT adversarial training pipeline, creating an effective and stable training procedure for the speech enhancement model. The resulting speech enhancement model, which we refer to as FINALLY, builds upon the HiFi++ architecture, augmented with a WavLM encoder and a novel training pipeline. Empirical results on various datasets confirm our model's ability to produce clear, high-quality speech at 48 kHz, achieving state-of-the-art performance in the field of speech enhancement. Demo page: https://samsunglabs.github.io/FINALLY-page

语音增强GAN高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。