用音乐结构分析增强扩散模型,生成更连贯的乐曲。
ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
- 结合舒恩克尔分析法与扩散模型,实现可解释的音乐生成。
- 提出分句融合方法,提升旋律与和声结构一致性。
- 支持用户控制生成过程,适合音乐创作与教学场景。
人工智能在音乐生成领域快速发展,现有符号化模型虽采用先进深度学习与扩散算法,但普遍存在结构性不连贯问题,尤其在和声-旋律结构上表现不足,且多为不可解释的“黑箱”。本文提出一种新框架ProGress(Prolongation-enhanced DiGress),融合舒恩克尔分析(SchA)与扩散建模思想,对Vignac等(2023)提出的DiGress模型进行改进,实现可解释、结构化的音乐生成。具体贡献包括:1)对DiGress模型的音乐生成适配;2)基于舒恩克尔分析的创新分句融合方法;3)支持用户调控生成过程的交互框架。人类评估结果显示,其性能优于现有最先进方法。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) for music generation is undergoing rapid developments, with recent symbolic models leveraging sophisticated deep learning and diffusion model algorithms. One drawback with existing models is that they lack structural cohesion, particularly on harmonic-melodic structure. Furthermore, such existing models are largely "black-box" in nature and are not musically interpretable. This paper addresses these limitations via a novel generative music framework that incorporates concepts of Schenkerian analysis (SchA) in concert with a diffusion modeling framework. This framework, which we call ProGress (Prolongation-enhanced DiGress), adapts state-of-the-art deep models for discrete diffusion (in particular, the DiGress model of Vignac et al., 2023) for interpretable and structured music generation. Concretely, our contributions include 1) novel adaptations of the DiGress model for music generation, 2) a novel SchA-inspired phrase fusion methodology, and 3) a framework allowing users to control various aspects of the generation process to create coherent musical compositions. Results from human experiments suggest superior performance to existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。