用人类作曲习惯训练模型,生成更富音乐性的符号化乐曲
MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
- 基于人类作曲的动机发展规则设计新模型
- 在POP909_M数据集上超越现有生成模型和ChatGPT-4
- 适合音乐创作、作曲辅助与人机协作场景
当前神经网络模型虽具强大序列预测能力,但其生成方式与人类作曲差异显著。作曲者通常从动机出发,依规则发展成完整作品,确保结构与变化规律。然而神经网络难以从数据中学习此类规则,导致生成音乐缺乏音乐性与多样性。本文提出将神经网络学习能力与人类作曲知识结合,构建首个标注动机及其变体的POP909_M数据集,并提出基于动机发展规则的文本到符号音乐生成模型MeloTrans。实验表明,MeloTrans在生成质量与多样性上均优于现有模型,甚至超越ChatGPT-4。结果凸显融合人类洞察与神经网络能力对提升符号音乐生成的重要性。
原文摘要 · Abstract (English)
At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$\_$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。