直接在频域操作,实现高质量语音修复与还原。
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
- 基于复数短时傅里叶变换系数进行生成式预训练。
- 多任务测试中均优于现有基线,目标说话人提取效果更优。
- 无需声码器,适合语音增强、修复等高保真场景。
本文提出一种用于高质量语音修复任务的生成式预训练基础模型。该模型直接在复数短时傅里叶变换系数上操作,无需任何声码器进行时域信号重建。因此,简化了合成流程,并消除了先前工作 SpeechFlow 中由梅尔谱图声码器带来的质量上限。所提方法在多个语音修复任务上进行了评估,包括语音去噪、带宽扩展、编解码器伪影去除和目标说话人提取。在所有场景中,微调预训练模型均显著优于强基线。尤其在目标说话人提取任务中,性能超越现有系统,包括使用 WavLM 等 SSL 预训练编码器的方法。代码与预训练检查点已在 NVIDIA NeMo 框架中公开。
原文摘要 · Abstract (English)
This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for time-domain signal reconstruction. As a result, our model simplifies the synthesis process and removes the quality upper-bound introduced by any mel-spectrogram vocoder compared to prior work SpeechFlow. The proposed method is evaluated on multiple speech restoration tasks, including speech denoising, bandwidth extension, codec artifact removal, and target speaker extraction. In all scenarios, finetuning our pretrained model results in superior performance over strong baselines. Notably, in the target speaker extraction task, our model outperforms existing systems, including those leveraging SSL-pretrained encoders like WavLM. The code and the pretrained checkpoints are publicly available in the NVIDIA NeMo framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。