arXiv:2510.00771eess.AScs.AI2025-10中稿 · ICASSP 2026被引 3

无需声码器的音频超分辨率方法,直接生成高质量音频波形。

UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching

  • 用流匹配模型直接建模复数谱系数分布,端到端生成波形。
  • 在多种上采样倍数下均达到48kHz高保真音频,性能领先。
  • 适合语音与通用音频增强,摆脱声码器性能限制。

本文提出一种无声码器的音频超分辨率框架,采用流匹配生成模型捕捉复数谱系数的条件分布,并通过逆短时傅里叶变换(iSTFT)直接重建波形,无需依赖外部声码器。与传统两阶段扩散模型先预测梅尔频谱再由预训练声码器合成波形的方法不同,该方法简化了端到端优化流程,克服了两阶段架构中最终音质受限于声码器性能的关键瓶颈。实验表明,该模型在多种上采样因子下均能稳定生成48kHz高保真音频,在语音和通用音频数据集上均达到当前最优性能。

原文摘要 · Abstract (English)

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage diffusion-based approaches that predict a mel-spectrogram and then rely on a pre-trained neural vocoder to synthesize waveforms, our method directly reconstructs waveforms via the inverse Short-Time Fourier Transform (iSTFT), thereby eliminating the dependence on a separate vocoder. This design not only simplifies end-to-end optimization but also overcomes a critical bottleneck of two-stage pipelines, where the final audio quality is fundamentally constrained by vocoder performance. Experiments show that our model consistently produces high-fidelity 48 kHz audio across diverse upsampling factors, achieving state-of-the-art performance on both speech and general audio datasets.

音频生成超分辨率流匹配无声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。