用语义与声学双重引导,实现跨场景音频超分辨率重建
SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
- 基于DiT架构,结合文本与频谱滚降嵌入进行条件控制
- 可将4~32kHz任意采样率音频稳定提升至44.1kHz,性能领先
- 适用于语音、音乐、音效等多类音频,输出语义对齐且高频一致
通用音频超分辨率旨在从低分辨率音频中预测高频频段成分,覆盖语音、音乐和音效等多种场景。现有基于扩散模型的方法常导致输出语义不一致,且难以保持高频重建一致性。本文提出SAGA-SR,一种融合语义与声学引导的通用音频超分辨率模型。该模型基于流匹配训练的DiT骨干网络,以文本和频谱滚降嵌入作为条件输入。得益于有效的条件引导,SAGA-SR可稳健地将4~32kHz任意采样率音频上采样至44.1kHz。客观与主观评估均表明,其在所有测试案例中均达到当前最优性能。相关音频示例与代码已公开。
原文摘要 · Abstract (English)
Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce semantically aligned outputs and struggle with consistent high-frequency reconstruction. In this paper, we propose SAGA-SR, a versatile audio SR model that combines semantic and acoustic guidance. Based on a DiT backbone trained with a flow matching objective, SAGA-SR is conditioned on text and spectral roll-off embeddings. Due to the effective guidance provided by its conditioning, SAGA-SR robustly upsamples audio from arbitrary input sampling rates between 4 kHz and 32 kHz to 44.1 kHz. Both objective and subjective evaluations show that SAGA-SR achieves state-of-the-art performance across all test cases. Sound examples and code for the proposed model are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。