TangoFlux 3.7秒生成30秒高保真音频,用新方法提升文本到音频对齐效果。
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
- 用流匹配+CLAP评分迭代优化偏好数据,解决音频生成缺乏标准答案难题。
- 515M参数模型在单张A40上3.7秒生成30秒44.1kHz音频,速度领先。
- 开源全部代码模型,适合研究高效文本到音频生成的开发者。
我们提出TangoFlux,一个拥有515M参数的高效文本到音频(TTA)生成模型,可在单张A40 GPU上仅用3.7秒生成长达30秒、采样率44.1kHz的音频。TTA模型对齐的一大难点在于难以构建偏好数据对,因缺乏类似大语言模型(LLMs)的可验证奖励或黄金标准答案。为此,我们提出CLAP-Ranked Preference Optimization(CRPO),一种通过迭代生成与优化偏好数据来提升对齐性能的新框架。实验表明,使用CRPO生成的音频偏好数据优于现有方案。基于该框架,TangoFlux在客观与主观评估中均达到当前最优表现。我们已开源全部代码与模型,以支持后续TTA生成研究。
原文摘要 · Abstract (English)
We introduce TangoFlux, an efficient Text-to-Audio (TTA) generative model with 515M parameters, capable of generating up to 30 seconds of 44.1kHz audio in just 3.7 seconds on a single A40 GPU. A key challenge in aligning TTA models lies in the difficulty of creating preference pairs, as TTA lacks structured mechanisms like verifiable rewards or gold-standard answers available for Large Language Models (LLMs). To address this, we propose CLAP-Ranked Preference Optimization (CRPO), a novel framework that iteratively generates and optimizes preference data to enhance TTA alignment. We demonstrate that the audio preference dataset generated using CRPO outperforms existing alternatives. With this framework, TangoFlux achieves state-of-the-art performance across both objective and subjective benchmarks. We open source all code and models to support further research in TTA generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。