arXiv:2508.13786cs.SDcs.AI2025-08被引 1

用动态事件图指导扩散模型,实现精准可控的文本生成音频

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

  • 将文本描述中的事件构建成动态图,融合语义与时间信息
  • 在多个数据集上超越现有方法,兼顾生成质量与控制精度
  • 适合需要精细控制音频内容和时序结构的研究者

可控文本到音频生成旨在根据文本描述合成音频,并满足用户指定的事件类型、时间顺序以及起始和结束时间戳,从而精确控制生成音频的内容与时间结构。尽管近期取得进展,现有方法仍面临准确时间定位、开放词汇可扩展性与实际效率之间的固有权衡。为此,我们提出DegDiT,一种面向开放词汇可控音频生成的动态事件图引导扩散变压器框架。DegDiT将描述中的事件编码为结构化动态图,图中节点设计用于表示语义特征、时间属性及事件间连接关系。通过图变压器整合这些节点,生成上下文感知的事件嵌入,作为扩散模型的引导信号。为确保高质量且多样化的训练数据,我们引入质量平衡的数据选择流程,结合层次化事件标注与多标准质量评分,构建出具有语义多样性的精炼数据集。此外,我们提出共识偏好优化机制,通过多个奖励信号的共识推动音频生成。在AudioCondition、DESED和AudioTime数据集上的大量实验表明,DegDiT在多种客观与主观评估指标上均达到最先进性能。

原文摘要 · Abstract (English)

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control over both the content and temporal structure of the generated audio. Despite recent progress, existing methods still face inherent trade-offs among accurate temporal localization, open-vocabulary scalability, and practical efficiency. To address these challenges, we propose DegDiT, a novel dynamic event graph-guided diffusion transformer framework for open-vocabulary controllable audio generation. DegDiT encodes the events in the description as structured dynamic graphs. The nodes in each graph are designed to represent three aspects: semantic features, temporal attributes, and inter-event connections. A graph transformer is employed to integrate these nodes and produce contextualized event embeddings that serve as guidance for the diffusion model. To ensure high-quality and diverse training data, we introduce a quality-balanced data selection pipeline that combines hierarchical event annotation with multi-criteria quality scoring, resulting in a curated dataset with semantic diversity. Furthermore, we present consensus preference optimization, facilitating audio generation through consensus among multiple reward signals. Extensive experiments on AudioCondition, DESED, and AudioTime datasets demonstrate that DegDiT achieves state-of-the-art performances across a variety of objective and subjective evaluation metrics.

音频生成扩散模型可控生成事件图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。