高效生成逼真音效,基于优化扩散Transformer
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
- 采用优化的扩散Transformer结构,提升收敛速度与效率
- 通过无分类器引导重缩放,高引导值下仍保音质与提示契合度
- 利用大模型生成合成描述词,增强预训练效果,适合音效生成研究者
我们提出EzAudio,一个文本到音频(T2A)生成框架,用于生成高质量、自然的声音效果。核心设计包括:(1) 提出EzAudio-DiT,一种针对音频潜在表示优化的扩散Transformer(DiT),提升收敛速度,并在参数量与内存使用上更高效;(2) 应用无分类器引导(CFG)重缩放技术,缓解高CFG得分下的保真度损失,提升提示遵循性而不牺牲音频质量;(3) 提出一种基于近期音频理解与大语言模型进展的合成描述词生成策略,用于增强T2A预训练。实验表明,EzAudio凭借计算高效架构与快速收敛,在客观与主观评估中均表现优异,提供高度真实的听觉体验。代码、数据与预训练模型已公开:https://haidog-yaqub.github.io/EzAudio-Page/
原文摘要 · Abstract (English)
We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio latent representations, improving convergence speed, as well as parameter and memory efficiency. (2) We apply a classifier-free guidance (CFG) rescaling technique to mitigate fidelity loss at higher CFG scores and enhancing prompt adherence without compromising audio quality. (3) We propose a synthetic caption generation strategy leveraging recent advances in audio understanding and LLMs to enhance T2A pretraining. We show that EzAudio, with its computationally efficient architecture and fast convergence, is a competitive open-source model that excels in both objective and subjective evaluations by delivering highly realistic listening experiences. Code, data, and pre-trained models are released at: https://haidog-yaqub.github.io/EzAudio-Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。