arXiv:2505.08175cs.SDcs.AI2025-05被引 19

通过对抗性后训练加速文本转音频,实现毫秒级生成。

Fast Text-to-Audio Generation with Adversarial Post-Training

  • 提出ARC后训练方法,结合相对对抗与对比判别器提升生成效率。
  • 在H100上75ms生成12秒44.1kHz立体声音频,移动端7秒完成。
  • 无需知识蒸馏,适合实时创作与边缘设备部署。

文本到音频系统虽性能不断提升,但推理速度慢,难以满足多数创意应用的延迟需求。本文提出对抗性相对对比(ARC)后训练方法,是首个不基于知识蒸馏的扩散/流模型对抗加速算法。该方法将近期提出的相对对抗框架扩展至扩散/流模型后训练,并引入新颖的对比判别器目标,以增强对提示词的遵循能力。结合对Stable Audio Open的多项优化,构建的模型可在H100上以约75ms生成约12秒、44.1kHz立体声音频,移动端约7秒完成,为目前最快文本到音频模型。

原文摘要 · Abstract (English)

Text-to-audio systems, while increasingly performant, are slow at inference time, thus making their latency unpractical for many creative applications. We present Adversarial Relativistic-Contrastive (ARC) post-training, the first adversarial acceleration algorithm for diffusion/flow models not based on distillation. While past adversarial post-training methods have struggled to compare against their expensive distillation counterparts, ARC post-training is a simple procedure that (1) extends a recent relativistic adversarial formulation to diffusion/flow post-training and (2) combines it with a novel contrastive discriminator objective to encourage better prompt adherence. We pair ARC post-training with a number optimizations to Stable Audio Open and build a model capable of generating $\approx$12s of 44.1kHz stereo audio in $\approx$75ms on an H100, and $\approx$7s on a mobile edge-device, the fastest text-to-audio model to our knowledge.

文本生成音频生成加速扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。