arXiv:2505.22106cs.SDcs.AI2025-05被引 7

用预训练模型加速文本转音频生成,10步即可超快出高质量音频。

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

  • 利用预训练模型生成确定性噪声对,学习一阶常微分方程路径。
  • 仅需10步采样即超越现有方法,推理速度比同类模型快3倍。
  • 适合追求高速高质音频生成的开发者与应用研究者。

扩散模型显著提升了音频生成的质量与多样性,但推理速度慢是其主要瓶颈。修正流通过学习直线常微分方程(ODE)路径提升推理速度,但需从头训练流匹配模型,且在低采样步数下表现不佳。为克服修正流局限并利用先进预训练扩散模型的优势,本文提出AudioTurbo,通过预训练文本到音频(TTA)模型生成确定性噪声对,学习一阶ODE路径。在AudioCaps数据集上的实验表明,该模型仅需10步采样即超越现有方法,推理速度较基于流匹配的加速模型降低至3步。

原文摘要 · Abstract (English)

Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.

文本转音频扩散模型加速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。