MeanAudio仅需一次函数计算即可生成高质量音频,速度比现有方法快100倍。
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- 采用均值流目标与引导速度,实现单步生成
- 在RTX 3090上实现实时因子0.013,速度提升100倍
- 适合需要快速生成音频的创作者和实时应用
近年来,文本到音频生成(TTA)取得了显著进展,为声音创作者提供了将灵感转化为生动音频的强大工具。然而,现有系统普遍存在推理速度慢的问题,严重影响创作效率与流畅性。本文提出MeanAudio,一种快速且忠实的文本到音频生成模型,可在单次函数评估(1-NFE)下生成逼真声音。该模型通过:(i) 均值流目标结合引导速度目标,显著加速推理;(ii) 改进的Flux风格变压器与双文本编码器,提升语义对齐与合成质量;(iii) 高效的瞬时-均值课程学习策略,加快收敛并支持在消费级显卡上训练。全面评估表明,MeanAudio在单步音频生成中达到顶尖性能:在单块NVIDIA RTX 3090上实现实时因子(RTF)0.013,较现有扩散模型系统提速100倍。同时在多步生成中表现优异,支持合成步骤间的平滑过渡。
原文摘要 · Abstract (English)
Recent years have witnessed remarkable progress in Text-to-Audio Generation (TTA), providing sound creators with powerful tools to transform inspirations into vivid audio. Yet despite these advances, current TTA systems often suffer from slow inference speed, which greatly hinders the efficiency and smoothness of audio creation. In this paper, we present MeanAudio, a fast and faithful text-to-audio generator capable of rendering realistic sound with only one function evaluation (1-NFE). MeanAudio leverages: (i) the MeanFlow objective with guided velocity target that significantly accelerates inference speed, (ii) an enhanced Flux-style transformer with dual text encoders for better semantic alignment and synthesis quality, and (iii) an efficient instantaneous-to-mean curriculum that speeds up convergence and enables training on consumer-grade GPUs. Through a comprehensive evaluation study, we demonstrate that MeanAudio achieves state-of-the-art performance in single-step audio generation. Specifically, it achieves a real-time factor (RTF) of 0.013 on a single NVIDIA RTX 3090, yielding a 100x speedup over SOTA diffusion-based TTA systems. Moreover, MeanAudio also shows strong performance in multi-step generation, enabling smooth transitions across successive synthesis steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。