用一步扩散蒸馏实现快速音频超分辨率,速度提升22倍
FlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation
- 通过扩散蒸馏实现单步生成,替代传统多步采样
- 在48kHz音频重建上达到当前最佳性能,推理速度提升22倍
- 适合需要高速音频修复的应用场景,如实时语音增强
通用音频超分辨率(SR)旨在从4kHz至32kHz的低采样率音频中恢复高频成分,适用于音乐、语音和音效等多种场景。现有基于扩散模型的方法因需大量采样步骤而推理缓慢。本文提出FlashSR,一种单步扩散模型,用于生成48kHz高保真音频。该方法通过三种目标实现扩散蒸馏:蒸馏损失、对抗损失与分布匹配蒸馏损失,显著加速推理。此外,我们设计了专为梅尔频谱图上的SR模型优化的SR Vocoder。在客观与主观评估中,FlashSR性能媲美当前最先进模型,同时推理速度提升约22倍。
原文摘要 · Abstract (English)
Versatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various domains such as music, speech, and sound effects. Previous diffusion-based SR methods suffer from slow inference due to the need for a large number of sampling steps. In this paper, we introduce FlashSR, a single-step diffusion model for versatile audio super-resolution aimed at producing 48kHz audio. FlashSR achieves fast inference by utilizing diffusion distillation with three objectives: distillation loss, adversarial loss, and distribution-matching distillation loss. We further enhance performance by proposing the SR Vocoder, which is specifically designed for SR models operating on mel-spectrograms. FlashSR demonstrates competitive performance with the current state-of-the-art model in both objective and subjective evaluations while being approximately 22 times faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。