arXiv:2604.09188cs.SD2026-04被引 1

用潜在空间流匹配提升音频超分辨率,更好还原高音细节。

LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching

  • 在潜在空间用条件流匹配生成高分辨率音频表示
  • 对语音、音效、音乐均表现优异,高频重建更精准
  • 适合需要高质量音频重建的研究与应用

音频超分辨率旨在从带宽受限的低分辨率音频中恢复缺失的高频细节,从而提升重构信号的自然度和听觉质量。然而,现有方法多直接在波形或时频域操作,不仅面临高维生成空间问题,且主要局限于语音任务,对音效和音乐等复杂音频类型改进有限。为此,我们提出LatentFlowSR,一种基于潜在表示空间的音频超分辨率新方法,利用条件流匹配(CFM)实现高效建模。首先训练一个抗噪鲁棒的自编码器,将低分辨率音频编码至连续潜在空间;随后,在低分辨率潜在表示条件下,通过一步常微分方程求解器,由高斯先验逐步生成对应高分辨率潜在表示;最后,使用预训练自编码器解码得到高分辨率音频。实验表明,该方法在多种音频类型与超分辨率设置下性能优于或媲美基线模型,验证了其出色的高频重建能力与泛化性能,为潜在空间建模在音频超分辨率中的有效性提供了有力证据。相关代码将在论文审稿完成后公开。

原文摘要 · Abstract (English)

Audio super-resolution aims to recover missing high-frequency details from bandwidth-limited low-resolution audio, thereby improving the naturalness and perceptual quality of the reconstructed signal. However, most existing methods directly operate in the waveform or time-frequency domain, which not only involves high-dimensional generation spaces but is also largely limited to speech tasks, leaving substantial room for improvement on more complex audio types such as sound effects and music. To mitigate these limitations, we introduce LatentFlowSR, a new audio super-resolution approach that leverages conditional flow matching (CFM) within a latent representation space. Specifically, we first train a noise-robust autoencoder, which encodes low-resolution audio into a continuous latent space. Conditioned on the low-resolution latent representation, a CFM mechanism progressively generates the corresponding high-resolution latent representation from a Gaussian prior with a one-step ordinary differential equation (ODE) solver. The resulting high-resolution latent representation is then decoded by the pretrained autoencoder to reconstruct the high-resolution audio. Experimental results demonstrate that LatentFlowSR achieves competitive or superior performance compared with baseline methods across various audio types and super-resolution settings. These results indicate that the proposed method possesses strong high-frequency reconstruction capability and robust generalization performance, providing compelling evidence for the effectiveness of latent-space modeling in audio super-resolution. All relevant code will be made publicly available upon completion of the paper review process.

音频重建流匹配潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。