用单步流匹配实现高效高质音频超分辨率。
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
- 引入流匹配模型,实现单步采样生成高质量音频。
- 在VCTK数据集上达到当前最佳的音质指标表现。
- 适合需要低延迟音频生成的应用场景。
音频超分辨率因其病态性而极具挑战性。近期扩散模型在该任务中展现出良好效果,但其需大量采样步骤,导致生成高保真音频时延迟过高。本文提出FLowHigh,将高效的流匹配生成模型应用于音频超分辨率,并设计专用于该任务的概率路径,有效捕捉高分辨率音频分布,显著提升重建质量。该方法仅通过单步采样即可生成高保真、高分辨率音频,适用于多种输入采样率。在VCTK基准数据集上的实验表明,FLowHigh在对数谱距离和ViSQOL指标上均达到当前最优性能,同时保持极低的计算开销。
原文摘要 · Abstract (English)
Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling steps, which causes significantly increased latency when synthesizing high-quality audio samples. In this paper, we propose FLowHigh, a novel approach that integrates flow matching, a highly efficient generative model, into audio super-resolution. We also explore probability paths specially tailored for audio super-resolution, which effectively capture high-resolution audio distributions, thereby enhancing reconstruction quality. The proposed method generates high-fidelity, high-resolution audio through a single-step sampling process across various input sampling rates. The experimental results on the VCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-art performance in audio super-resolution, as evaluated by log-spectral distance and ViSQOL while maintaining computational efficiency with only a single-step sampling process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。