arXiv:2601.01294cs.SDcs.AI2026-01被引 5

无需训练,通过噪声注入与结构约束实现音乐音色迁移。

Diffusion Timbre Transfer Via Mutual Information Guided Inpainting

  • 在潜在空间中对乐器身份关键通道注入噪声。
  • 反向扩散早期阶段强制保留原曲旋律节奏结构。
  • 兼容文本/音频条件,适合音乐风格迁移场景。

我们将音色迁移视为音乐音频的推理时编辑问题。基于一个强大的预训练潜空间扩散模型,提出一种轻量级方法,无需额外训练:(i) 针对最能反映乐器身份的潜在通道进行逐维度噪声注入;(ii) 在反向扩散早期阶段引入结构钳制机制,重新施加输入的旋律与节奏结构。该方法直接作用于音频潜在表示,支持文本/音频条件(如 CLAP)。我们讨论了设计选择,分析了音色变化与结构保持之间的权衡,并证明简单的推理时控制即可有效引导预训练模型实现风格迁移任务。

原文摘要 · Abstract (English)

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise injection that targets latent channels most informative of instrument identity, and (ii) an early-step clamping mechanism that re-imposes the input's melodic and rhythmic structure during reverse diffusion. The method operates directly on audio latents and is compatible with text/audio conditioning (e.g., CLAP). We discuss design choices,analyze trade-offs between timbral change and structural preservation, and show that simple inference-time controls can meaningfully steer pre-trained models for style-transfer use cases.

音色迁移扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。