arXiv:2509.23727cs.SDcs.AI2025-09中稿 · ICME 2026被引 5

用混合引导提升音频生成质量,不需重训练。

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance

  • 融合多种引导信号与交互项,优化采样过程。
  • 在相同速度下,文本到音频生成优于单一引导方法。
  • 适合追求高质量音频生成的研究者与开发者。

基于扩散模型的音频生成系统设计已从数据空间、网络架构和条件技术等多个角度被广泛研究,但多数创新需重新训练模型。在采样阶段,无分类器引导(CFG)被普遍采用以增强生成质量,但常牺牲多样性,导致性能不佳。尽管近期自引导(AG)方法提供了保持多样性的新方向,但在音频生成中表现仍逊于CFG。本文提出AudioMoG,一种无需大量训练资源即可提升文本到音频(T2A)和视频到音频(V2A)生成质量的改进采样方法。我们分析了CFG与AG各自的优劣,基于洞察提出混合引导框架,整合多种引导信号及其交互项(如模型的无条件劣化版本),以最大化累积优势。实验表明,在相同推理速度下,该方法在不同采样步数上均一致优于单一路径引导,在T2A、V2A、文本到音乐及图像生成任务中均有显著提升。演示样本可访问:https://audiomog.github.io。

原文摘要 · Abstract (English)

The design of diffusion-based audio generation systems has been investigated from diverse perspectives, such as data space, network architecture, and conditioning techniques, while most of these innovations require model re-training. In sampling, classifier-free guidance (CFG) has been uniformly adopted to enhance generation quality by strengthening condition alignment. However, CFG often compromises diversity, resulting in suboptimal performance. Although the recent autoguidance (AG) method proposes another direction of guidance that maintains diversity, its direct application in audio generation has so far underperformed CFG. In this work, we introduce AudioMoG, an improved sampling method that enhances text-to-audio (T2A) and video-to-audio (V2A) generation quality without requiring extensive training resources. We start with an analysis of both CFG and AG, examining their respective advantages and limitations for guiding diffusion models. Building upon our insights, we introduce a mixture-of-guidance framework that integrates diverse guidance signals with their interaction terms (e.g., the unconditional bad version of the model) to maximize cumulative advantages. Experiments show that, given the same inference speed, our approach consistently outperforms single guidance in T2A generation across sampling steps, concurrently showing advantages in V2A, text-to-music, and image generation. Demo samples are available at: https://audiomog.github.io.

音频生成扩散模型引导机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。