arXiv:2603.04122cs.SDcs.LG2026-03被引 1

FastWave用轻量化扩散模型实现48kHz音频超分辨率,速度快且资源消耗低。

FastWave: Optimized Diffusion Model for Audio Super-Resolution

  • 基于优化的扩散模型框架,重构高采样率音频信号
  • 仅需50 GFLOPs算力与130万参数,训练速度显著提升
  • 适合追求高效推理的音频处理应用开发者

音频超分辨率旨在高质量重建给定信号,使其表现如同以更高采样率采集。现有方法包括扩散模型和流模型(较慢)、生成对抗网络(较快),但均依赖高参数量网络,导致训练与推理计算成本高昂。本文提出一种新方案,重新利用近期扩散模型训练技术,应用于任意采样率到48 kHz的超分辨率任务。所提模型FastWave在性能上优于NU-Wave 2,接近当前最优水平。其计算复杂度约为50 GFLOPs,参数量仅1.3 M,训练资源需求更低、速度更快,相比多数近期扩散与流模型解决方案有明显优势。代码已公开。

原文摘要 · Abstract (English)

Audio Super-Resolution is a set of techniques aimed at high-quality estimation of the given signal as if it would be sampled with higher sample rate. Among suggested methods there are diffusion and flow models (which are considered slower), generative adversarial networks (which are considered faster), however both approaches are currently presented by high-parametric networks, requiring high computational costs both for training and inference. We propose a solution to both these problems by re-considering the recent advances in the training of diffusion models and applying them to super-resolution from any to 48 kHz sample rate. Our approach shows better results than NU-Wave 2 and is comparable to state-of-the-art models. Our model called FastWave has around 50 GFLOPs of computational complexity and 1.3 M parameters and can be trained with less resources and significantly faster than the majority of recently proposed diffusion- and flow-based solutions. The code has been made publicly available.

音频超分辨率扩散模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。