arXiv:2605.04547cs.SDcs.AI2026-05

根据训练阶段动态调整音频生成策略,提升效率与质量。

Stage-adaptive audio diffusion modeling

  • 按训练进度自适应调整语义引导、采样策略和结构正则化。
  • 在文本生成与超分辨率任务中收敛更快,指标提升3%~5%。
  • 适合追求高效音频生成的科研与工业开发者使用。

基于扩散的音频生成与修复近年来显著提升了跨条件场景的表现,包括文本条件生成和音频条件超分辨率。然而,训练过程仍计算开销大,多数现有流程依赖静态优化方案,固定训练信号的重要性权重。本文指出,效率低下的主要原因是语义学习与生成精炼之间平衡随训练进程不断变化:初期侧重对齐条件的语义结构与全局组织,后期则更关注时间一致性、感知保真度与细节优化。为此,我们提出一种基于训练时SSL空间差异斜率的进度变量,表征语义进展。基于此信号,设计三种互补的阶段感知机制:早期使用的衰减式SSL引导、由阶段变量驱动的自适应时间步采样,以及参数空间聚类收敛后激活的结构感知正则化。在文本条件音频生成与音频条件超分辨率任务上验证,所提方法在收敛行为与主生成指标、频谱重建指标上均优于标准静态基线,提升3%-5%。结果表明,将外部引导、内部组织与优化重心视为阶段相关组件,而非固定配置,可显著提升音频扩散模型训练效率。

原文摘要 · Abstract (English)

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution. However, training audio diffusion models remains computationally expensive, and most existing pipelines still rely on static optimization recipes that treat the relative importance of training signals as fixed throughout learning. In this work, we argue that a major source of inefficiency lies in the evolving balance between semantic acquisition and generation-oriented refinement. Early training places stronger emphasis on acquiring condition-aligned semantic structure and coarse global organization, whereas later training increasingly emphasizes temporal consistency, perceptual fidelity, and fine-detail refinement. To characterize this evolving balance, we introduce a progress-based regime variable derived from the training-time slope of an SSL-space discrepancy, which measures semantic progress during training. Based on this signal, we develop three complementary stage-aware mechanisms: decayed SSL guidance for early semantic bootstrapping, self-adaptive timestep sampling driven by the regime variable, and structure-aware regularization activated from convergent grouped organization in parameter space. We evaluate these mechanisms on text-conditioned audio generation and audio-conditioned super-resolution. Across both settings, the proposed stage-aware strategies improve convergence behavior and yield gains on the primary generation and spectral reconstruction metrics over standard static baselines. These results support the view that efficient audio diffusion training can benefit from treating external guidance, internal organization, and optimization emphasis as stage-dependent components rather than fixed ingredients.

音频生成扩散模型自适应训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。