arXiv:2509.17609cs.SDcs.LG2025-09NeurIPS被引 17

用潜在空间桥梁模型实现高质量音频超分辨率,支持任意采样率到192kHz的无缝提升。

Audio Super-Resolution with Latent Bridge Models

  • 在潜在空间构建桥梁模型,直接从低分辨率潜码生成高分辨率潜码。
  • 在多个数据集上达到最优客观与主观质量,首次实现任意到192kHz的音频超分辨率。
  • 适用于语音、音乐和通用音频,适合音频后期制作与高保真重建场景。

音频超分辨率(SR)旨在将低分辨率(LR)波形上采样至高分辨率(HR)版本。现有方法多依赖扩散模型或桥梁模型,但常因生成先验信息不足导致质量不佳。本文提出潜在桥梁模型(Latent Bridge Models, LBMs),将音频波形压缩至连续潜在空间,设计潜码到潜码的生成流程,天然匹配LR到HR的上采样过程,充分挖掘LR波形中的指导性先验信息。为应对高质量样本稀缺问题,引入频率感知的LBMs,以输入先验与目标频率作为条件,使模型在训练阶段显式学习任意到任意的上采样能力。进一步设计级联式LBMs并提出两种先验增强策略,首次实现超过48kHz的音频上采样,支持平滑级联式超分辨率,显著提升音频后期处理灵活性。在VCTK、ESC-50、Song-Describer基准数据集及两个内部测试集上的全面实验表明,本方法在任意到48kHz的音频超分辨率任务中均达到当前最优的客观与主观质量,并创下任意到192kHz音频超分辨率的首项记录。

原文摘要 · Abstract (English)

Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, while previous methods often suffer from sub-optimal upsampling quality due to their uninformative generation prior. Towards high-quality audio super-resolution, we present a new system with latent bridge models (LBMs), where we compress the audio waveform into a continuous latent space and design an LBM to enable a latent-to-latent generation process that naturally matches the LR-toHR upsampling process, thereby fully exploiting the instructive prior information contained in the LR waveform. To further enhance the training results despite the limited availability of HR samples, we introduce frequency-aware LBMs, where the prior and target frequency are taken as model input, enabling LBMs to explicitly learn an any-to-any upsampling process at the training stage. Furthermore, we design cascaded LBMs and present two prior augmentation strategies, where we make the first attempt to unlock the audio upsampling beyond 48 kHz and empower a seamless cascaded SR process, providing higher flexibility for audio post-production. Comprehensive experimental results evaluated on the VCTK, ESC-50, Song-Describer benchmark datasets and two internal testsets demonstrate that we achieve state-of-the-art objective and perceptual quality for any-to-48kHz SR across speech, audio, and music signals, as well as setting the first record for any-to-192kHz audio SR. Demo at https://AudioLBM.github.io/.

音频超分潜在模型频率感知级联结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。