arXiv:2605.30594eess.AS2026-05

用单模型实现多带宽音频超分辨率,性能超越现有方法。

FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution

论文配图:FiPA-SR -- FiLM-Conditioned Perceptually Informed Audio Super-Resolution
图 1 · 摘自论文原文
  • 引入FiLM层按输入带宽动态调整重建过程。
  • 在8/20/32 kHz下均优于AudioSR,且推理速度提升60倍。
  • 仅需约1/3的显存,适合实时音频增强场景。

音频带宽扩展旨在从受限带宽信号中重建缺失的高频内容。本文提出FiPA-SR,一种基于GAN的感知架构,可在单一模型中处理不同输入带宽。在AEROMamba_P框架基础上,引入FiLM层以根据输入带宽自适应调整重建过程。在MUSDB数据集上的实验表明,FiPA-SR在8、20和32 kHz输入采样率下均优于当前最优的AudioSR模型。此外,该架构仅需约3倍更少的GPU内存,且推理速度比基于扩散模型的基线快60倍以上。

原文摘要 · Abstract (English)

Audio bandwidth extension aims to reconstruct missing high-frequency content from bandlimited signals. This paper proposes FiPA-SR, a GAN-based perceptual architecture capable of handling different input bandwidths within a single model. Building upon the previous $\textrm{AEROMamba}_\textrm{P}$ framework, the proposed model incorporates FiLM layers to adapt the reconstruction process according to the respective bandwidth. Experiments on the MUSDB dataset show that FiPA-SR outperforms the state-of-the-art AudioSR model across 8, 20, and 32 kHz input sampling rates. Moreover, the proposed architecture uses approximately 3$\times$ less GPU memory and performs inference more than 60$\times$ faster than the diffusion-based baseline.

音频修复GAN超分辨率高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。