arXiv:2502.14893cs.CVcs.AI2025-02NAACL被引 7

首个多模态乐谱理解数据集,让视觉大模型读懂乐谱图文。

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

  • 构建跨模态乐谱图像与ABC记谱法对齐的训练数据
  • 在100万+记录上训练出性能显著提升的NotaGPT模型
  • 适合音乐生成、智能作曲等方向的研究者使用

符号化音乐以两种形式呈现:二维直观的乐谱图像和一维标准的文本记谱序列。尽管大语言模型在音乐领域展现出巨大潜力,但现有研究主要集中在单模态文本序列上。通用视觉语言模型仍缺乏乐谱理解能力。针对这一空白,我们提出了NOTA——首个大规模综合性多模态乐谱理解数据集,包含1,019,237条记录,覆盖全球三个地区,涵盖三项任务。基于该数据集,我们训练了NotaGPT,一种音乐记谱视觉大语言模型。具体而言,先进行跨模态对齐预训练,使乐谱图像中的音符与其在ABC记谱法中的文本表示对齐;随后分阶段训练,依次完成基础音乐信息提取与乐谱分析任务。实验表明,NotaGPT-7B在音乐理解任务上取得显著提升,验证了NOTA数据集与训练流程的有效性。相关数据集已开源至https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset。

原文摘要 · Abstract (English)

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music, current research has primarily focused on unimodal symbol sequence text. Existing general-domain visual language models still lack the ability of music notation understanding. Recognizing this gap, we propose NOTA, the first large-scale comprehensive multimodal music notation dataset. It consists of 1,019,237 records, from 3 regions of the world, and contains 3 tasks. Based on the dataset, we trained NotaGPT, a music notation visual large language model. Specifically, we involve a pre-alignment training phase for cross-modal alignment between the musical notes depicted in music score images and their textual representation in ABC notation. Subsequent training phases focus on foundational music information extraction, followed by training on music notation analysis. Experimental results demonstrate that our NotaGPT-7B achieves significant improvement on music understanding, showcasing the effectiveness of NOTA and the training pipeline. Our datasets are open-sourced at https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset.

多模态乐谱理解视觉语言模型音乐AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。