用6万小时钢琴谱训练自回归模型,生成更连贯的音乐并实现顶尖分类性能。
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
- 基于6万小时乐谱预训练,用高质量子集微调生成音乐和分类任务。
- 音乐续写质量超越主流符号生成方法,接近商用音频生成模型。
- 对比学习提取的音乐嵌入在少量标注数据下即可高效适应下游任务。
我们研究了在大量符号化独奏钢琴乐谱上训练的生成式自回归Transformer模型的能力。首先在约6万小时的音乐数据上进行预训练,再使用一个较小但高质量的子集,微调模型以生成音乐续写、完成符号分类任务,并通过适配SimCLR框架生成通用对比性MIDI嵌入。在钢琴续写连贯性评估中,生成模型的表现优于主流符号生成技术,且与专有音频生成模型相当。在MIR分类基准测试中,冻结的对比模型表示在线性探测实验中达到当前最优水平;直接微调则证明了预训练表示的泛化能力,通常仅需数百个标注样本即可适配下游任务。
原文摘要 · Abstract (English)
We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。