arXiv:2409.10753eess.AScs.SD2024-09中稿 · ICASSP 2025被引 38

对比扩散模型训练目标,提升语音增强的听感质量。

Investigating Training Objectives for Generative Speech Enhancement

  • 基于分数模型与薛定谔桥框架对比训练机制。
  • 新设计感知损失使语音听感更自然,效果更优。
  • 适合语音增强与生成模型研究者参考。

生成式语音增强在噪声环境下提升语音质量方面取得显著进展。现有多种基于扩散的框架,采用不同训练目标与学习方法。本文聚焦于分数驱动的生成模型与薛定谔桥框架,通过一系列全面实验比较其性能与训练行为差异。进一步提出一种针对薛定谔桥框架定制的新型感知损失函数,显著提升增强语音的感知质量。所有实验代码与预训练模型均公开,以促进该领域研究发展。

原文摘要 · Abstract (English)

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrödinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrödinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain.

语音增强扩散模型感知损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。