融合预测与生成模型,实现单模型通用语音增强。
A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
- 设计双分支模型:预测分支直接还原语音,生成分支优化扩散去噪。
- 在多种数据集上优于现有方法,显著提升语音质量与可懂度。
- 适合需要高鲁棒性语音增强的场景,如嘈杂环境下的通信系统。
设计单一模型以抑制多种失真并提升语音质量,即通用语音增强(USE),具有广阔前景。相较于基于监督学习的预测方法,基于扩散的生成模型因具备从严重损坏信息中重建的能力而展现出更大潜力。然而,在极端恶劣条件下易引入伪影,且扩散模型推理步骤多,计算开销大。为协同利用预测与生成的优势并克服各自缺陷,本文提出名为PGUSE的通用语音增强模型,采用预测与生成建模相结合的策略。模型包含两个分支:预测分支直接从降质信号中预测干净样本,生成分支优化扩散模型的去噪目标。通过输出融合与截断扩散机制有效整合两者:前者直接结合双分支输出,后者利用预测分支的初始估计修改反向扩散过程。在多个数据集上的大量实验验证了该模型优于当前最优基线,证明了预测与生成建模互补带来的优势。
原文摘要 · Abstract (English)
It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning-based predictive methods, diffusion-based generative models have shown greater potential due to the generative capacities from degraded speech with severely damaged information. However, artifacts may be introduced in highly adverse conditions, and diffusion models often suffer from a heavy computational burden due to many steps for inference. In order to jointly leverage the superiority of prediction and generation and overcome the respective defects, in this work we propose a universal speech enhancement model called PGUSE by combining predictive and generative modeling. Our model consists of two branches: the predictive branch directly predicts clean samples from degraded signals, while the generative branch optimizes the denoising objective of diffusion models. We utilize the output fusion and truncated diffusion scheme to effectively integrate predictive and generative modeling, where the former directly combines results from both branches and the latter modifies the reverse diffusion process with initial estimates from the predictive branch. Extensive experiments on several datasets verify the superiority of the proposed model over state-of-the-art baselines, demonstrating the complementarity and benefits of combining predictive and generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。