arXiv:2409.12560eess.AScs.SD2024-09中稿 · ICASSP 2025被引 13

仅用自然语言描述实现精细音频生成,提升可控性与质量

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

  • 用自然语言同时指定内容和风格,无需额外帧级条件
  • 基于流模型的扩散变换器,生成速度更快且效果更优
  • 自动生成细粒度标注数据,缓解训练数据稀缺问题

当前文本到音频(TTA)模型主要依赖粗粒度文本输入,限制了对音频内容与风格的精细控制。部分研究尝试通过引入帧级条件或控制网络提升粒度,但导致系统复杂且需参考帧级信息。为此,我们提出 AudioComposer,一种仅依赖自然语言描述(NLDs)即可同时提供内容与风格控制的新框架。采用基于流的扩散变换器结合交叉注意力机制,有效将文本信息融入音频生成过程,既能同步处理内容与风格,又比其他架构加速生成。此外,我们设计了一种新颖的自动数据模拟流水线,构建具有细粒度文本描述的数据集,显著缓解该领域数据稀缺问题。实验表明,仅使用 NLDs 输入时,该框架在生成质量和可控性上均超越现有 SOTA 模型,且模型规模更小。

原文摘要 · Abstract (English)

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.

音频生成文本到音频自然语言控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。