arXiv:2605.16681eess.AScs.SD2026-05综述被引 6

从判别模型到生成模型,系统梳理音频超分辨率发展脉络

A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models

论文配图:A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
图 1 · 摘自论文原文
  • 对比判别与生成模型在音频重建中的设计差异
  • 实验证明生成模型在保真度与感知质量上更优
  • 适合音频处理、生成模型研究者参考

音频超分辨率(SR),又称带宽扩展(BWE),旨在从低分辨率或带限信号中重建高保真信号,由于高频内容缺失导致的不确定性,该任务本质上是病态的。本文综述了该领域的进展,重点关注从判别性映射到现代生成建模的范式转变。首先回顾早期基于深度神经网络(DNN)的判别模型,将BWE/SR视为确定性映射问题,易出现回归均值效应和频谱过平滑。随后系统分析生成方法,包括自回归(AR)模型、变分自编码器(VAEs)、生成对抗网络(GANs)、扩散与基于得分模型、基于流的方法以及薛定谔桥。我们考察了关键设计要素:表示域、架构、条件机制,以及重建保真度、感知质量、鲁棒性与计算效率之间的权衡。通过在代表性判别与生成方法上进行统一实验,提供了受控的实证证据。此外,讨论了大语言模型(LLMs)与多模态基础模型等新兴方向,并指出感知评估、实际部署与真实世界泛化方面的开放挑战。本文通过结构化分类与统一视角,建立全面基础,为从确定性点估计迈向分布感知生成建模提供实用路线图。

原文摘要 · Abstract (English)

Audio super-resolution (SR), also referred to as bandwidth extension (BWE), aims to reconstruct high-fidelity signals from low-resolution (LR) or band-limited (BL) observations, an inherently ill-posed task due to the ambiguity of missing high-frequency (HF) content. This survey provides a comprehensive overview of the field, with a particular focus on the paradigm shift from discriminative mapping to modern generative modeling. We first review early discriminative deep neural network (DNN) models, which formulate BWE/SR as a deterministic mapping problem and are prone to regression-to-the-mean effects and spectral over-smoothing. We then systematically review generative approaches, including autoregressive (AR) models, variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion and score-based models, flow-based methods, and Schrödinger bridges. Across these approaches, we examine key design aspects, including representation domain, architecture, conditioning mechanisms, and trade-offs among reconstruction fidelity, perceptual quality, robustness, and computational efficiency. We further conduct unified experiments on representative discriminative and generative methods to provide controlled empirical evidence for these trade-offs. Furthermore, we discuss emerging directions involving large language models (LLMs) and multimodal foundation models, and highlight open challenges in perceptual evaluation, practical deployment, and real-world generalization. By providing a structured taxonomy and unified perspective, this survey establishes a comprehensive foundation and offers a practical roadmap for advancing BWE/SR from deterministic point estimation toward distribution-aware generative modeling.

音频生成生成模型超分辨率语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。