arXiv:2409.15321eess.AScs.SD2024-09中稿 · MLSP 2024被引 8

用扩散模型实现多乐器音色迁移,一模型支持多种配对。

WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion

  • 基于双边去噪扩散模型,端到端完成音色转换。
  • 支持混合音频与单个乐器间的音色迁移,44.1kHz可用。
  • 单模型兼容多乐器配对,无需为每对重训练。

随着基于扩散的生成模型日益普及,研究者正积极探索其在音乐合成与风格变换等领域的应用。本文关注音色迁移任务——在保留核心音乐元素的同时,无缝改变乐曲的乐器特性。我们提出WaveTransfer,一种用于音色迁移的端到端扩散模型。具体采用双边去噪扩散模型(BDDM)进行噪声调度搜索。该模型可对音频混音及单个乐器进行音色迁移。尤为突出的是,它具备多类型音色迁移的通用性,可在同一模型中支持多种独特乐器配对,无需为每对单独训练。此外,不同于近期工作仅限于16 kHz,WaveTransfer可在多种采样率下训练,包括音乐领域标准的44.1 kHz,对音乐社区具有重要意义。

原文摘要 · Abstract (English)

As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.

音色迁移扩散模型音乐生成多乐器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。