一个模型搞定多采样率音频生成,还能用低采样数据提升高保真音质。
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
- 将采样率作为生成条件,统一控制扩散模型的音频输出
- 低采样率预训练可显著提升高采样率音频质量
- 适合需要跨采样率部署的音频生成应用场景
我们提出SRC-gAudio,一种新型音频生成模型,可在单一模型架构中实现跨多种采样率的文本到音频生成。该模型将采样率作为生成条件,引导基于扩散的音频生成过程。实验表明,该模型能有效生成受控采样率的音频。此外,我们发现使用大规模低采样率数据进行预训练,可显著提升各类指标下的高采样率音频生成质量。
原文摘要 · Abstract (English)
We introduce SRC-gAudio, a novel audio generation model designed to facilitate text-to-audio generation across a wide range of sampling rates within a single model architecture. SRC-gAudio incorporates the sampling rate as part of the generation condition to guide the diffusion-based audio generation process. Our model enables the generation of audio at multiple sampling rates with a single unified model. Furthermore, we explore the potential benefits of large-scale, low-sampling-rate data in enhancing the generation quality of high-sampling-rate audio. Through extensive experiments, we demonstrate that SRC-gAudio effectively generates audio under controlled sampling rates. Additionally, our results indicate that pre-training on low-sampling-rate data can lead to significant improvements in audio quality across various metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。