arXiv:2608.05222cs.SDeess.AS2026-08

用扩散模型生成高质量长时长音乐,文本控制更精准。

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

论文配图:Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
图 1 · 摘自论文原文
  • 基于自回归扩散模型生成音乐,提升时长与连贯性。
  • 在19,345个文本模板上训练,显著提升生成质量与多样性。
  • 适合音乐创作新手和专业作曲者使用。

文本控制的符号化音乐生成因灵活直观而受到关注,但以往方法在质量、多样性、可控性和时长方面存在局限。本文提出Diff-Symbo,一种基于潜在扩散模型(LDM)的创新方法,可生成高质量、多样且长时长的符号化音乐。为解决文本-音乐数据集缺失问题,我们利用大语言模型构建了包含19,345个文本模板的综合性数据集。同时设计音乐信息编码器,在降低训练开销的同时提取更有效的控制表征。结合自回归策略,该方法显著提升了生成音乐的时长与作曲一致性。实验表明,相比GPT-4、MuseCoco和Multitrack Music Transformer(MMT)等基线模型,Diff-Symbo在文本可控性、生成时长与音乐质量上均有显著提升。作为该领域的先驱模型,Diff-Symbo为基于LDM的可控高质量音乐创作提供了重要贡献。

原文摘要 · Abstract (English)

Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous approaches tend to generate symbolic music with compromising quality, diversity, controllability and limited duration. In this paper, we present Diff-Symbo, an innovative method that uses latent diffusion model (LDM) to generate high-quality, diverse and long-duration symbolic music. To address the lack of text-symbolic music dataset, we develop a comprehensive dataset with 19,345 text templates by employing large language model. Furthermore, we design a music information encoder to reduce the training overhead while extracting more effective control representations. Given textual descriptions, our proposed method leverages LDM to improve the quality and diversity of music generation. Our method also improves the duration and the compositional consistency of music generation through an autoregressive approach. Experimental results show significant improvements of Diff-Symbo in text controllability, duration, and the quality of generated music compared to the baseline models such as GPT-4, MuseCoco and Multitrack Music Transformer (MMT). As one of the pioneer models in this field, Diff-Symbo paves the way towards controllable and high-quality symbolic music composition based on LDM, offering valuable contributions to both music amateurs and practitioners.

音乐生成扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。