arXiv:2604.10628cs.SDcs.CL2026-04被引 1

构建首个巴洛克时期乐谱的专家标注莉莉朋德数据集,提升音乐理解效果。

BMdataset: A Musicologically Curated LilyPond Dataset

论文配图:BMdataset: A Musicologically Curated LilyPond Dataset
图 1 · 摘自论文原文
  • 用专家从原手稿转录的莉莉朋德代码构建高质量音乐数据集
  • 小规模专业数据集微调效果优于大规模噪声数据预训练
  • 适合研究符号化音乐表示学习与古典音乐分析的研究者

符号化音乐研究长期依赖MIDI数据集,而基于文本的排版格式如莉莉朋德尚未被充分探索。本文提出BMdataset,一个由专家从巴洛克时期手稿直接转录的莉莉朋德乐谱数据集,包含393个作品、2,646个乐章,附有作曲家、体裁、乐器配置和段落属性等元数据。基于此,我们开发了LilyBERT,一种通过扩展115个莉莉朋德特定词元并进行掩码语言建模预训练的CodeBERT变体。在跨域的Mutopia语料上线性探测显示,尽管其规模较小(约9000万词元),仅在BMdataset上微调的效果已优于在完整PDMX语料(约150亿词元)上的连续预训练,用于作曲家和风格分类任务。结合广泛预训练与领域特定微调可获得最佳性能(作曲家识别准确率达84.3%),证实两类数据互补。我们公开数据集、分词器和模型,为莉莉朋德上的表示学习建立基准。

原文摘要 · Abstract (English)

Symbolic music research has relied almost exclusively on MIDI-based datasets; text-based engraving formats such as LilyPond remain unexplored for music understanding. We present BMdataset, a musicologically curated dataset of 393 LilyPond scores (2,646 movements) transcribed by experts directly from original Baroque manuscripts, with metadata covering composer, musical form, instrumentation, and sectional attributes. Building on this resource, we introduce LilyBERT (weights can be found at https://huggingface.co/csc-unipd/lilybert), a CodeBERT-based encoder adapted to symbolic music through vocabulary extension with 115 LilyPond-specific tokens and masked language model pre-training. Linear probing on the out-of-domain Mutopia corpus shows that, despite its modest size (~90M tokens), fine-tuning on BMdataset alone outperforms continuous pre-training on the full PDMX corpus (~15B tokens) for both composer and style classification, demonstrating that small, expertly curated datasets can be more effective than large, noisy corpora for music understanding. Combining broad pre-training with domain-specific fine-tuning yields the best results overall (84.3% composer accuracy), confirming that the two data regimes are complementary. We release the dataset, tokenizer, and model to establish a baseline for representation learning on LilyPond.

符号音乐音乐生成深度学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。