arXiv:2409.06126eess.AScs.SD2024-09被引 2

先降噪再用扩散模型修复语音,提升音质与细节。

VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion

  • 两阶段设计:先降噪,再用扩散模型基于说话人特征和内容信息修复语音。
  • 修复后语音在带宽、混响消除和缺失段补全上均有明显改善。
  • 适合对语音质量要求高、但可容忍轻微清晰度下降的场景。

降噪算法虽能改善语音质量,但过度降噪会损害目标语音,降低可懂性和质量。本文提出一种基于语音转换(VC)的显式语音恢复方法,在降噪后通过扩散模型实现高质量语音重建,该模型以目标说话人嵌入和从去噪语音中提取的内容信息为条件。此修复过程可实现带宽扩展、去混响及语音补全等增强效果。实验表明,该两阶段NS+VC框架在客观指标上优于单阶段增强模型,语音质量更高,但可懂性略低。为此,本文进一步提出一种内容编码器适应方法,提升噪声环境下内容提取的鲁棒性。

原文摘要 · Abstract (English)

Noise suppression (NS) algorithms are effective in improving speech quality in many cases. However, aggressive noise suppression can damage the target speech, reducing both speech intelligibility and quality despite removing the noise. This study proposes an explicit speech restoration method using a voice conversion (VC) technique for restoration after noise suppression. We observed that high-quality speech can be restored through a diffusion-based voice conversion stage, conditioned on the target speaker embedding and speech content information extracted from the de-noised speech. This speech restoration can achieve enhancement effects such as bandwidth extension, de-reverberation, and in-painting. Our experimental results demonstrate that this two-stage NS+VC framework outperforms single-stage enhancement models in terms of output speech quality, as measured by objective metrics, while scoring slightly lower in speech intelligibility. To further improve the intelligibility of the combined system, we propose a content encoder adaptation method for robust content extraction in noisy conditions.

语音修复扩散模型降噪语音转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。