首个专为乐谱设计的基础视觉模型,可高效理解音乐符号结构。
MuSViT: A Foundation Vision Model for Sheet Music Representation

- 用掩码自编码预训练970万页乐谱,分两阶段提升模型泛化能力。
- 在4个下游任务中表现优于现有方法,线性探测下仍领先于通用视觉模型。
- 首次证明乐谱表示能直接捕捉音乐符号结构,适合音乐智能研究者。
基础模型已彻底改变视觉与语言处理,提供可跨任务复用的丰富表征。但作为音乐语言的视觉编码,乐谱缺乏此类领域专用骨干。本文提出MuSViT(Music Score Vision Transformer):首个针对乐谱表征的视觉变换器基础模型,采用掩码自编码在970万页IMSLP乐谱上进行预训练。为应对真实乐谱复杂性,采用两阶段课程学习:先在排版合成乐谱上预热,再在完整IMSLP语料库上大规模训练。在四个下游任务上评估:全页与谱线级乐谱识别、音乐符号检测、乐谱难度分类,涵盖线性探测(冻结编码器)与微调两种场景。线性探测下,MuSViT持续优于现代视觉编码器,表明通用表征即使规模庞大,仍系统性地不足于捕捉乐谱的结构化符号特性。微调时,MuSViT普遍超越特定任务的最先进方法。额外的嵌入-转录一致性分析显示,MuSViT的表示空间直接编码了符号化音乐结构——而其他编码器的嵌入与乐谱内容无相关性。这些结果确立了MuSViT作为乐谱理解的基础骨干。
原文摘要 · Abstract (English)
Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet music, as a visual encoding of musical language, lacks such a strong domain-specific backbone. We introduce MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation -- a ViT encoder pre-trained via Masked Autoencoders on 9.7 million pages from the IMSLP. To handle the complexity of real-world scores, we adopt a two-stage curriculum: a synthetic warm-up on typeset scores followed by large-scale training on the full IMSLP corpus. We evaluate MuSViT on four downstream tasks -- full-page and staff-level music score recognition, music symbol detection, and score difficulty classification -- under two scenarios: linear probing (frozen encoder) and fine-tuning. Under linear probing, MuSViT consistently outperforms modern vision encoders, revealing that general-purpose representations, regardless of scale, fall systematically short on the structured symbolic properties of musical notation. Under fine-tuning, MuSViT generally improves upon task-specific state-of-the-art methods. An additional embedding-transcription consistency analysis reveals that MuSViT encodes symbolic musical structure directly in its representation space -- unlike other encoders, whose embeddings do not correlate with music notation content. These results establish MuSViT as a foundation backbone for sheet music understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。