用扩散模型预热初始化,一套方法搞定多种音频转换任务。
Audio-to-Audio via Diffusion Warm Initialization

- 用预训练扩散模型的中间状态做初始化,无需额外训练
- 选对初始化时间可让音色迁移等任务保持输入忠实度与目标分布匹配
- 无需加噪,引导信号本身就能当初始状态,适合快速部署
本文提出扩散预热初始化,一种简单而有效的音频到音频转换通用方法。我们将其应用于音色迁移、MIDI转真实音频合成及多项音频增强任务。通过对音色迁移进行详尽的实证分析,研究了初始化时间 $t_ ext{init}$ 的作用,采用基于音高的杰卡德距离和弗雷切特音频距离评估输入保真度与目标分布对齐程度。结果为 $t_ ext{init}$ 的选择提供了实用指导,表明一旦合理设定,单一预训练扩散模型结合预热初始化即可支持多类转换任务,无需特定任务训练或条件控制。尽管方法简单,性能仍优于许多专为这些任务设计的复杂流水线。此外发现,预热初始化未必需要显式加噪,引导信号本身常可作为反向扩散过程的有效初始状态。整体表明,该方法构成复杂音频转换流水线的基础模块。
原文摘要 · Abstract (English)
In this paper, we propose diffusion warm initialization as a simple yet effective approach for a range of audio-to-audio transformation tasks. To illustrate the generality of the approach, we demonstrate its use in timbre transfer, MIDI-to-Real synthesis, and multiple audio enhancement tasks. We conduct a detailed empirical analysis on timbre transfer to investigate the role of the initialization time $t_\text{init}$. The effect of $t_\text{init}$ is evaluated using pitch-based Jaccard Distance and Fréchet Audio Distance to quantify faithfulness to the input signal and alignment with the target distribution. Our results provide practical guidance for selecting $t_\text{init}$ and show that, once properly chosen, a single pretrained diffusion model combined with warm initialization can support multiple transformation objectives without task-specific training or conditioning. Despite its simplicity, this approach already achieves competitive results when compared with more complex pipelines designed specifically for these tasks. We further observe that warm initialization does not necessarily require explicit noise injection, as the guide signal itself can often serve as a valid initialization state for the backward diffusion process. Together, these findings show that warm initialization provides a simple and effective framework that serves as a fundamental building block for more complex audio transformation pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。