arXiv:2606.16435eess.AScs.SD2026-06

一个模型搞定语音生成与编辑,省去专用模块。

Unified Audio Generation and Editing via Joint Condition Modeling and Progressive Training

论文配图:Unified Audio Generation and Editing via Joint Condition Modeling and Progressive Training
图 1 · 摘自论文原文
  • 用联合条件建模和分层位置编码统一处理文本生成与音频编辑
  • 在6个编辑任务上表现接近专用模型,且无性能下降
  • 适合需要多功能音频系统的研究者与开发者

随着多媒体应用对音频需求的增长,众多先进的音频生成方法相继出现。现有研究通常将文本到音频(TTA)及其他相关任务(如基于指令的音频编辑)视为独立问题,采用特定任务的架构或模块。这种缺乏统一建模范式的方式显著增加了系统构建的开销与复杂度,也限制了可扩展性。为此,我们提出 AudioWeave,一个无需额外任务专用组件的统一模型,同时支持 TTA 和音频编辑。具体地,我们设计了一种联合条件建模方法,结合因子化位置嵌入,使扩散变换器主干网络能够处理异构输入(包括 TTA 与音频编辑)。此外,我们提出渐进式多阶段训练策略,缓解多任务间的竞争与灾难性遗忘问题,从而维持各任务性能,甚至在某些方面有所提升。在 TTA 任务及六个音频编辑任务上的实验结果表明,该统一模型性能与专用模型相当,为统一音频生成模型的进一步探索奠定了基础。

原文摘要 · Abstract (English)

With the growing focus on audio in multimedia applications, numerous advanced works on audio generation have emerged. Existing studies typically treat text-to-audio (TTA) and other related audio generation tasks, such as instruction-based audio editing, as independent challenges, adopting task-specific architectures or modules. This absence of a unified modeling paradigm substantially increases the overhead and complexity of building a system for both audio generation and editing, while also leading to limited scalability. To address this issue, we introduce AudioWeave, a unified model for TTA and audio editing without additional task-specific components. Specifically, we propose a joint condition modeling approach with a factorized position embedding, enabling the diffusion transformer backbone to operate under heterogeneous inputs of TTA and audio editing. We further propose a progressive multistage training strategy to mitigate task competition and catastrophic forgetting caused by interference among multiple tasks. This in turn helps maintain the performance of each individual task and may even lead to improvements in certain aspects. Experimental results on TTA task and six audio editing tasks show that our unified model achieves competitive performance with task-specific models, laying a groundwork for further exploration of unified audio generation models.

音频生成扩散模型统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。