用大模型训练范式生成高质量古典乐谱,提升音乐美感与可控性。
NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
- 基于预训练+微调+强化学习的LLM训练范式生成乐谱
- 在1.6百万首曲子上预训练,9千首名作上微调
- 无需人工标注即可提升生成质量,适合音乐创作与研究
我们提出NotaGen,一种符号化音乐生成模型,旨在探索生成高质量古典乐谱的潜力。受大语言模型成功启发,NotaGen采用预训练、微调和强化学习范式。它在160万首以ABC记谱法表示的音乐作品上进行预训练,并在约9,000首高质量古典作品上,基于“时期-作曲家-配器”提示进行微调。针对强化学习,我们提出CLaMP-DPO方法,在无需人类标注或预定义奖励的情况下进一步提升生成质量与可控性。实验表明,CLaMP-DPO在不同架构和编码方案的模型中均有效。主观A/B测试显示,NotaGen生成结果优于基线模型,接近人类创作水平,显著提升了符号化音乐生成的音乐审美表现。
原文摘要 · Abstract (English)
We introduce NotaGen, a symbolic music generation model aiming to explore the potential of producing high-quality classical sheet music. Inspired by the success of Large Language Models (LLMs), NotaGen adopts pre-training, fine-tuning, and reinforcement learning paradigms (henceforth referred to as the LLM training paradigms). It is pre-trained on 1.6M pieces of music in ABC notation, and then fine-tuned on approximately 9K high-quality classical compositions conditioned on "period-composer-instrumentation" prompts. For reinforcement learning, we propose the CLaMP-DPO method, which further enhances generation quality and controllability without requiring human annotations or predefined rewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolic music generation models with different architectures and encoding schemes. Furthermore, subjective A/B tests show that NotaGen outperforms baseline models against human compositions, greatly advancing musical aesthetics in symbolic music generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。