arXiv:2409.12346cs.SDeess.AS2024-09被引 16

一个模型同时实现音乐分离与多轨生成,还能自动生成编曲。

Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models

  • 用潜空间扩散模型联合建模多轨音乐的共享音乐上下文。
  • 在Slakh2100数据集上,分离与生成效果均优于现有方法。
  • 适合音乐生成、分离及编曲自动化研究者使用。

扩散模型在音乐生成和音源分离任务中展现出强大潜力。尽管仍处于早期阶段,但将两者整合到统一框架的趋势正在形成,因为二者都涉及生成符合音乐逻辑的片段,可视为同一生成过程的不同方面。本文提出一种基于潜空间扩散的多轨生成模型,通过学习共享音乐语境下各音轨的联合概率分布,实现音源分离与多轨音乐合成的同步完成,并支持在已知部分音轨的情况下生成其余音轨的编曲。模型在Slakh2100数据集上训练,与现有同时生成与分离模型相比,在客观评价指标上于音源分离、音乐生成及编曲生成任务中均取得显著提升。音频示例见 https://msg-ld.github.io/。

原文摘要 · Abstract (English)

Diffusion models have recently shown strong potential in both music generation and music source separation tasks. Although in early stages, a trend is emerging towards integrating these tasks into a single framework, as both involve generating musically aligned parts and can be seen as facets of the same generative process. In this work, we introduce a latent diffusion-based multi-track generation model capable of both source separation and multi-track music synthesis by learning the joint probability distribution of tracks sharing a musical context. Our model also enables arrangement generation by creating any subset of tracks given the others. We trained our model on the Slakh2100 dataset, compared it with an existing simultaneous generation and separation model, and observed significant improvements across objective metrics for source separation, music, and arrangement generation tasks. Sound examples are available at https://msg-ld.github.io/.

音乐生成扩散模型音源分离多轨合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。