arXiv:2409.08514cs.SDcs.AI2024-09中稿 · ICASSP 2025, Demo …被引 14

Apollo通过频段建模提升高采样率音频修复质量。

Apollo: Band-sequence Modeling for High-Quality Audio Restoration

  • 分频段建模,显式处理高低频关系以提升重建一致性。
  • 在MUSDB18-HQ和MoisesDB上优于现有SR-GAN模型,尤其在多乐器混合场景。
  • 适合需要高质量音频还原的研究者与音视频处理工程师。

音频修复在现代社会愈发重要,不仅因高端播放设备对高品质听觉体验的需求,也因生成式音频模型的发展要求高保真音频。通常,音频修复任务是从受损输入中预测无失真的音频,常采用GAN框架以平衡感知质量与失真度。由于音频退化主要集中在中高频段(尤其是编码器导致),核心挑战在于设计一个既能保留低频信息,又能精准重建高质量中高频内容的生成器。受近期高采样率音乐分离、语音增强及音频编码模型启发,我们提出Apollo,一种专为高采样率音频修复设计的生成模型。Apollo引入显式的频段分割模块,建模不同频段间的关系,从而实现更连贯、更高品质的修复效果。在MUSDB18-HQ和MoisesDB数据集上的评估表明,Apollo在多种比特率和音乐类型下持续优于现有SR-GAN模型,尤其在包含多个乐器与人声的复杂场景中表现突出。该模型显著提升音乐修复质量的同时保持计算效率。Apollo源码已公开于https://github.com/JusperLee/Apollo。

原文摘要 · Abstract (English)

Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio models necessitate high-fidelity audio. Typically, audio restoration is defined as a task of predicting undistorted audio from damaged input, often trained using a GAN framework to balance perception and distortion. Since audio degradation is primarily concentrated in mid- and high-frequency ranges, especially due to codecs, a key challenge lies in designing a generator capable of preserving low-frequency information while accurately reconstructing high-quality mid- and high-frequency content. Inspired by recent advancements in high-sample-rate music separation, speech enhancement, and audio codec models, we propose Apollo, a generative model designed for high-sample-rate audio restoration. Apollo employs an explicit frequency band split module to model the relationships between different frequency bands, allowing for more coherent and higher-quality restored audio. Evaluated on the MUSDB18-HQ and MoisesDB datasets, Apollo consistently outperforms existing SR-GAN models across various bit rates and music genres, particularly excelling in complex scenarios involving mixtures of multiple instruments and vocals. Apollo significantly improves music restoration quality while maintaining computational efficiency. The source code for Apollo is publicly available at https://github.com/JusperLee/Apollo.

音频修复生成模型频段建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。