arXiv:2505.15254cs.SDeess.AS2025-05中稿 · INTERSPEECH 2025被引 1

用扩散模型实现无参考的语音修复与声线转换,提升降噪后语音质量。

Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework

  • 两阶段设计:先生成式降噪,再声线转换增强语音
  • 在多个数据集上达到与顶尖方法相当的客观评分
  • 适合语音修复、录音降噪及语音合成领域研究者

我们提出一种结合无说话人依赖语音修复与声线转换(VC)的语音增强系统,以获得接近录音室级别的语音质量。尽管声线转换模型通常用于改变说话人特征,但在目标说话人与源说话人相同时,也可用于语音修复。然而,由于VC模型对噪声敏感,我们在系统前端引入生成式语音修复(GSR)模型,该模型在无需目标说话人信息的情况下实现降噪并恢复受损语音。随后,VC阶段利用干净的说话人嵌入进行引导,进一步优化输出语音。通过这种两阶段方法,我们在多个数据集上的语音质量客观指标得分达到了与当前最先进(SOTA)方法相当的水平。

原文摘要 · Abstract (English)

We propose a speech enhancement system that combines speaker-agnostic speech restoration with voice conversion (VC) to obtain a studio-level quality speech signal. While voice conversion models are typically used to change speaker characteristics, they can also serve as a means of speech restoration when the target speaker is the same as the source speaker. However, since VC models are vulnerable to noisy conditions, we have included a generative speech restoration (GSR) model at the front end of our proposed system. The GSR model performs noise suppression and restores speech damage incurred during that process without knowledge about the target speaker. The VC stage then uses guidance from clean speaker embeddings to further restore the output speech. By employing this two-stage approach, we have achieved speech quality objective metric scores comparable to state-of-the-art (SOTA) methods across multiple datasets.

语音修复扩散模型声线转换降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。