arXiv:2606.31259cs.SDcs.AI2026-06

仅用文本描述训练一键生成音频模型,速度更快且无需配对音频数据

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

论文配图:SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
图 1 · 摘自论文原文
  • 通过文本无音频的蒸馏方式,从预训练模型学习生成先验
  • 仅用4.5万条文本就能训练出效果接近多步模型的一键生成系统
  • 适合追求高效推理、缺乏标注音频数据的研究者与开发者

基于扩散模型的文本到音频(TTA)生成在合成质量上表现优异,但因迭代多步去噪导致推理延迟高。现有的一步生成方法虽缓解此问题,仍需成对的文本-音频数据进行蒸馏。为此,我们提出SwiftAudio,一种仅使用文本描述即可完成音频无关蒸馏的一步式TTA框架。具体地,我们将变分得分蒸馏(VSD)适配至音频领域,并引入时间平滑正则化目标以促进连贯的潜在音频表示。该设计使学生模型可在无需配对音频监督的情况下继承教师模型的生成先验,支持仅约4.5万条文本的高效训练。在AudioCaps和Clotho数据集上的实验表明,SwiftAudio在严格一步方法中达到当前最佳性能,显著缩小了与多步扩散系统之间的差距。

原文摘要 · Abstract (English)

Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step approaches alleviate this issue but still rely on paired text--audio data during distillation. To address these limitations, we propose SwiftAudio, a one-step TTA framework that performs audio-free distillation from a pretrained diffusion teacher using only text captions. Specifically, we adapt Variational Score Distillation (VSD) to the audio domain and introduce a temporal smoothness regularization objective to encourage coherent latent audio representations. This design enables the student model to inherit the teacher's generative prior without requiring paired audio supervision and allows effective training with only approximately 45K captions. Experiments on AudioCaps and Clotho demonstrate that SwiftAudio achieves state-of-the-art performance among strict one-step methods and substantially narrows the gap to multi-step diffusion systems. Project page: https://swiftaudio.org/

文本到音频扩散模型蒸馏高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。