arXiv:2409.08155eess.AS2024-09被引 4

用图神经网络生成有节奏和长程结构的中文流行音乐

Hierarchical Symbolic Pop Music Generation with Graph Neural Networks

  • 构建多图模型表示节奏与乐句结构
  • 两阶段生成:先生成4小节乐句,再构建全曲结构
  • 能捕捉和弦、音高分布及乐句特征

音乐具有复杂的内在结构,将其表示为图可有效捕捉多层次关系。尽管已有多种深度生成技术用于音乐创作,但基于图的音乐生成研究仍较少。早期工作仅限于旋律生成,近期的多声部音乐生成方法未能考虑长期结构。本文提出一种多图方法,用于表征中文流行音乐的节奏模式与乐句结构。为此,我们设计了两步生成流程:首先在MIDI数据集上训练变分自编码器以生成4小节乐句;其次在歌曲结构标签上训练另一变分自编码器以生成完整曲式。实验表明,模型能学习训练集中大部分结构性细节,包括和弦与音高频率分布,以及乐句属性。

原文摘要 · Abstract (English)

Music is inherently made up of complex structures, and representing them as graphs helps to capture multiple levels of relationships. While music generation has been explored using various deep generation techniques, research on graph-related music generation is sparse. Earlier graph-based music generation worked only on generating melodies, and recent works to generate polyphonic music do not account for longer-term structure. In this paper, we explore a multi-graph approach to represent both the rhythmic patterns and phrase structure of Chinese pop music. Consequently, we propose a two-step approach that aims to generate polyphonic music with coherent rhythm and long-term structure. We train two Variational Auto-Encoder networks - one on a MIDI dataset to generate 4-bar phrases, and another on song structure labels to generate full song structure. Our work shows that the models are able to learn most of the structural nuances in the training dataset, including chord and pitch frequency distributions, and phrase attributes.

音乐生成图神经网络流行音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。