用判别模型的潜在表示增强生成模型,提升语音分离质量
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
- 将判别模型的隐含特征注入生成扩散模型,融合两者优势
- 在MUSDB数据集上,语音信噪比提升3.7%,音质分离度提升10.0%
- 适合关注语音增强中生成与判别协同的科研人员和工程师
基于扩散的生成模型因其对复杂语音数据分布的建模能力,在语音与人声增强任务中取得了显著成果。尽管这些模型在未见声学环境下的泛化性能良好,但其保真度仍不及针对特定声学条件训练的判别模型。本文提出Ex-Diff,一种新的基于得分的扩散模型,通过整合判别模型产生的潜在表示,以提升语音与人声增强效果,融合生成与判别模型的优势。在广泛使用的MUSDB数据集上的实验表明,相较于基线扩散模型,该方法在语音增强任务中分别实现了3.7%的SI-SDR相对提升和10.0%的SI-SIR相对提升。此外,案例研究进一步揭示了生成与判别模型在此场景下的互补性。
原文摘要 · Abstract (English)
Diffusion-based generative models have recently achieved remarkable results in speech and vocal enhancement due to their ability to model complex speech data distributions. While these models generalize well to unseen acoustic environments, they may not achieve the same level of fidelity as the discriminative models specifically trained to enhance particular acoustic conditions. In this paper, we propose Ex-Diff, a novel score-based diffusion model that integrates the latent representations produced by a discriminative model to improve speech and vocal enhancement, which combines the strengths of both generative and discriminative models. Experimental results on the widely used MUSDB dataset show relative improvements of 3.7% in SI-SDR and 10.0% in SI-SIR compared to the baseline diffusion model for speech and vocal enhancement tasks, respectively. Additionally, case studies are provided to further illustrate and analyze the complementary nature of generative and discriminative models in this context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。