用大规模通用音乐预训练,提升对巴赫等四位作曲家风格的精准生成能力。
From Generality to Mastery: Composer-Style Symbolic Music Generation via Large-Scale Pre-training
- 先在流行、民谣、古典音乐上预训练,再用小样本作曲家数据微调。
- 在巴赫、莫扎特等四人作品上生成效果更贴近原作风格,音乐性更优。
- 适合研究作曲家风格建模与小样本音乐生成的学者与创作者。
尽管可控符号化音乐生成取得进展,特定控制模态仍面临数据稀缺问题。作曲家风格生成是典型例子,每位作照者仅有少量乐曲可用来建模风格与基本音乐元素(如旋律、和弦、节奏)。本文探讨如何利用广泛语料中学习的通用音乐知识,增强对特定作曲家风格的掌握能力,聚焦钢琴作品生成。方法采用两阶段训练:首先在包含流行、民谣、古典音乐的大规模语料上预训练基于REMI的音乐生成模型;随后在四个著名作曲家(巴赫、莫扎特、贝多芬、肖邦)的少量人工验证数据集上微调,使用轻量级适配器模块以风格标识作为条件。通过客观与主观评估,结果表明该方法优于消融实验与基线模型,在风格准确性和音乐美感上均有提升。同时,我们观察到模型如何从通用预训练中构建音乐概念,并通过专精微调优化风格理解。
原文摘要 · Abstract (English)
Despite progress in controllable symbolic music generation, data scarcity remains a challenge for certain control modalities. Composer-style music generation is a prime example, as only a few pieces per composer are available, limiting the modeling of both styles and fundamental music elements (e.g., melody, chord, rhythm). In this paper, we investigate how general music knowledge learned from a broad corpus can enhance the mastery of specific composer styles, with a focus on piano piece generation. Our approach follows a two-stage training paradigm. First, we pre-train a REMI-based music generation model on a large corpus of pop, folk, and classical music. Then, we fine-tune it on a small, human-verified dataset from four renowned composers, namely Bach, Mozart, Beethoven, and Chopin, using a lightweight adapter module to condition the model on style indicators. To evaluate the effectiveness of our approach, we conduct both objective and subjective evaluations on style accuracy and musicality. Experimental results demonstrate that our method outperforms ablations and baselines, achieving more precise composer-style modeling and better musical aesthetics. Additionally, we provide observations on how the model builds music concepts from the generality pre-training and refines its stylistic understanding through the mastery fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。