arXiv:2510.08878cs.SDcs.AI2025-10ACL被引 11

通过渐进扩散模型实现精准语音控制的文本生成音频

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

  • 分步建模文本、时间、音素信息,提升生成可控性
  • 在客观与主观评测中显著优于现有方法
  • 适合需要精确时序和清晰语音的应用场景

近期研究探索了具有细粒度控制信号(如精确时序控制或可理解语音内容)的文本到音频(TTA)生成。然而,受限于数据稀缺,其大规模生成性能仍不理想。本文将可控TTA生成重新定义为多任务学习问题,提出一种渐进式扩散建模方法ControlAudio。该方法通过分步策略,有效拟合依赖于更细粒度信息(包括文本、时间、音素特征)的分布。首先,提出涵盖标注与模拟的数据构建方法,增强文本-时间-音素序列中的条件信息;其次,在模型训练阶段,先在大规模文本-音频对上预训练扩散变压器(DiT),实现可扩展的TTA生成,再逐步融合时间与音素特征,采用统一语义表示以扩展可控性;最后,在推理阶段,提出渐进式引导生成,按顺序强调更细粒度信息,契合DiT固有的粗到精采样特性。大量实验表明,ControlAudio在时序准确性与语音清晰度方面均达到当前最优水平,客观与主观评估均显著优于现有方法。演示样本见:https://control-audio.github.io/Control-Audio。

原文摘要 · Abstract (English)

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale is still compromised. In this study, we recast controllable TTA generation as a multi-task learning problem and introduce a progressive diffusion modeling approach, ControlAudio. Our method adeptly fits distributions conditioned on more fine-grained information, including text, timing, and phoneme features, through a step-by-step strategy. First, we propose a data construction method spanning both annotation and simulation, augmenting condition information in the sequence of text, timing, and phoneme. Second, at the model training stage, we pretrain a diffusion transformer (DiT) on large-scale text-audio pairs, achieving scalable TTA generation, and then incrementally integrate the timing and phoneme features with unified semantic representations, expanding controllability. Finally, at the inference stage, we propose progressively guided generation, which sequentially emphasizes more fine-grained information, aligning inherently with the coarse-to-fine sampling nature of DiT. Extensive experiments show that ControlAudio achieves state-of-the-art performance in terms of temporal accuracy and speech clarity, significantly outperforming existing methods on both objective and subjective evaluations. Demo samples are available at: https://control-audio.github.io/Control-Audio.

音频生成扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。