对比扩散模型训练目标,提升语音增强的听感质量。
Investigating Training Objectives for Generative Speech Enhancement
- 基于分数模型与薛定谔桥框架对比训练机制。
- 新设计感知损失使语音听感更自然,效果更优。
- 适合语音增强与生成模型研究者参考。
生成式语音增强在噪声环境下提升语音质量方面取得显著进展。现有多种基于扩散的框架,采用不同训练目标与学习方法。本文聚焦于分数驱动的生成模型与薛定谔桥框架,通过一系列全面实验比较其性能与训练行为差异。进一步提出一种针对薛定谔桥框架定制的新型感知损失函数,显著提升增强语音的感知质量。所有实验代码与预训练模型均公开,以促进该领域研究发展。
原文摘要 · Abstract (English)
Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrödinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrödinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。