arXiv:2409.02845cs.SDcs.MM2024-09被引 8

让音乐生成更像作曲:多轨协同控制,自动配器更精准。

Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

  • 基于潜空间扩散模型,联合建模多轨音乐的上下文关系。
  • 支持全量生成与部分补全(如补钢琴轨),在多个指标上优于现有方法。
  • 适合需要精细多轨控制的音乐创作、智能作曲系统开发者。

扩散模型在跨模态音频生成任务中表现优异,如文本到音乐生成。然而,现有文本控制的音乐生成模型通常仅关注风格、情绪等全局属性,难以满足复杂多层音乐编排的需求。音乐创作需对各乐器轨在节奏、动态、和声与旋律上精确协调,而文本提示难以提供足够控制力。为此,本文将MusicLDM扩展为多轨生成模型,通过学习共享上下文下各轨的联合概率分布,实现多轨音乐的条件或无条件生成。模型还能进行编排生成,即在给定部分轨道(如鼓和贝斯)的情况下生成其余轨道(如钢琴)。实验表明,在总生成和编排生成任务中,本模型在客观指标上均显著优于现有方法。

原文摘要 · Abstract (English)

Diffusion models have shown promising results in cross-modal generation tasks involving audio and music, such as text-to-sound and text-to-music generation. These text-controlled music generation models typically focus on generating music by capturing global musical attributes like genre and mood. However, music composition is a complex, multilayered task that often involves musical arrangement as an integral part of the process. This process involves composing each instrument to align with existing ones in terms of beat, dynamics, harmony, and melody, requiring greater precision and control over tracks than text prompts usually provide. In this work, we address these challenges by extending the MusicLDM, a latent diffusion model for music, into a multi-track generative model. By learning the joint probability of tracks sharing a context, our model is capable of generating music across several tracks that correspond well to each other, either conditionally or unconditionally. Additionally, our model is capable of arrangement generation, where the model can generate any subset of tracks given the others (e.g., generating a piano track complementing given bass and drum tracks). We compared our model with an existing multi-track generative model and demonstrated that our model achieves considerable improvements across objective metrics for both total and arrangement generation tasks.

音乐生成扩散模型多轨协同编排生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。