用潜在扩散模型实现48kHz高保真语音修复,解决多畸变还原难题。
High-Resolution Speech Restoration with Latent Diffusion Model
- 基于潜在扩散架构,统一处理多种语音畸变。
- 在高频谐波恢复上超越现有方法,生成更自然的语音细节。
- 适合专业级语音修复,兼顾听感与计算效率。
传统语音增强方法常仅针对单一类型畸变,而现有生成模型在处理多种畸变时,往往难以重建音素与高频谐波,导致呼吸、喘息等伪影,降低语音可懂度。这些模型还计算开销大,多数仅限于宽带频段输出,限制了其在专业场景的应用。为此,我们提出Hi-ResLDM,一种基于潜在扩散的新型生成模型,可同时去除多种畸变,将语音恢复至48kHz的录音室级质量。我们在多个指标上对比最先进的基于GAN与条件流匹配(CFM)的方法,结果表明,Hi-ResLDM在高频带细节重建方面表现更优。该模型不仅在非侵入性指标上领先,且在人工评估中持续获得偏好,在侵入性评估中也具备竞争力,适用于高分辨率语音修复任务。
原文摘要 · Abstract (English)
Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that reduce the intelligibility of reconstructed speech. These models are also computationally demanding, and many solutions are restricted to producing outputs in the wide-band frequency range, which limits their suitability for professional applications. To address these challenges, we propose Hi-ResLDM, a novel generative model based on latent diffusion designed to remove multiple distortions and restore speech recordings to studio quality, sampled at 48kHz. We benchmark Hi-ResLDM against state-of-the-art methods that leverage GAN and Conditional Flow Matching (CFM) components, demonstrating superior performance in regenerating high-frequency-band details. Hi-ResLDM not only excels in non-instrusive metrics but is also consistently preferred in human evaluation and performs competitively on intrusive evaluations, making it ideal for high-resolution speech restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。