arXiv:2501.00794eess.AScs.SD2025-01被引 3

用自监督训练的流匹配模型,一键修复语音中的噪音、混响等问题。

VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration

  • 基于条件流匹配与无分类器引导,无需成对数据即可修复语音。
  • 在合成数据上训练,可统一处理噪声、混响、带宽限制等多种退化。
  • 适用于短语和长对话,修复效果显著且泛化能力强。

我们提出 VoiceRestore,一种利用自监督训练的流匹配变换器,用于恢复语音录音质量。该方法针对短时与长时语音中常见的多种退化问题(如背景噪声、混响、压缩伪影和带宽限制)提供统一解决方案。通过条件流匹配与无分类器引导,模型在无需成对清洁与退化数据的情况下,学习将劣质语音映射为高质量录音。文中详细描述了训练流程、条件流匹配框架及模型架构,并验证其在真实语音任务中的泛化能力,涵盖短语及长篇独白或对话。定性与定量评估表明,该方法在不同长度与退化类型下均具高效灵活性。

原文摘要 · Abstract (English)

We present VoiceRestore, a novel approach to restoring the quality of speech recordings using flow-matching Transformers trained in a self-supervised manner on synthetic data. Our method tackles a wide range of degradations frequently found in both short and long-form speech recordings, including background noise, reverberation, compression artifacts, and bandwidth limitations - all within a single, unified model. Leveraging conditional flow matching and classifier free guidance, the model learns to map degraded speech to high quality recordings without requiring paired clean and degraded datasets. We describe the training process, the conditional flow matching framework, and the model's architecture. We also demonstrate the model's generalization to real-world speech restoration tasks, including both short utterances and extended monologues or dialogues. Qualitative and quantitative evaluations show that our approach provides a flexible and effective solution for enhancing the quality of speech recordings across varying lengths and degradation types.

语音修复流匹配自监督Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。