arXiv:2412.19351cs.SDcs.CL2024-12ICML被引 22

系统探索文本生成音频的模型设计空间,提出更优生成方案

ETTA: Elucidating the Design Space of Text-to-Audio Models

  • 构建大规模合成音效数据集AF-Synthetic,提升训练质量
  • 对比多种架构与采样策略,发现生成速度与质量的权衡关系
  • 模型在复杂创意描述生成上表现优异,适合内容创作场景

近年来,文本到音频(TTA)合成取得显著进展,使用户能通过自然语言提示生成合成音频以丰富创作流程。然而,数据、模型架构、训练目标函数和采样策略对基准任务的影响尚未充分理解。为全面解析TTA模型的设计空间,我们开展大规模实证实验,聚焦扩散模型与流匹配模型。贡献包括:1)构建高质量合成标注数据集AF-Synthetic,由音频理解模型生成;2)系统比较不同架构、训练与推理设计选择;3)分析采样方法及其在生成质量与推理速度间的帕累托曲线。基于该分析,我们提出最优模型Elucidated Text-To-Audio(ETTA)。在AudioCaps与MusicCaps上,ETTA优于公开数据训练的基线,且媲美专有数据训练模型。此外,其在复杂、富有想象力的文本生成音频任务中表现更佳,挑战了现有基准。

原文摘要 · Abstract (English)

Recent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the effects of data, model architecture, training objective functions, and sampling strategies on target benchmarks are not well understood. With the purpose of providing a holistic understanding of the design space of TTA models, we set up a large-scale empirical experiment focused on diffusion and flow matching models. Our contributions include: 1) AF-Synthetic, a large dataset of high quality synthetic captions obtained from an audio understanding model; 2) a systematic comparison of different architectural, training, and inference design choices for TTA models; 3) an analysis of sampling methods and their Pareto curves with respect to generation quality and inference speed. We leverage the knowledge obtained from this extensive analysis to propose our best model dubbed Elucidated Text-To-Audio (ETTA). When evaluated on AudioCaps and MusicCaps, ETTA provides improvements over the baselines trained on publicly available data, while being competitive with models trained on proprietary data. Finally, we show ETTA's improved ability to generate creative audio following complex and imaginative captions -- a task that is more challenging than current benchmarks.

文本生成音频扩散模型音频合成设计空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。