arXiv:2605.18436cs.CV2026-05

构建首个真实历史手写乐谱数据集,助力机器识别珍贵音乐遗产。

A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation

论文配图:A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation
图 1 · 摘自论文原文
  • 收集1309页真实历史手写乐谱,含MusicXML转录与符号标注
  • 是目前最大且最真实的西方记谱法手写乐谱数据集
  • 适合训练和评估端到端及目标检测类乐谱识别系统

大量音乐文化遗产已被图书馆、博物馆和档案馆数字化。然而,尽管深度学习取得进展,光学乐谱识别(OMR)仍难以将这些乐谱变为机器可读格式,主要原因是缺乏在真实条件下训练系统的数据集。为此,MusiCorpus数据集提供了1,309页历史乐谱,以手写为主,包含MusicXML转录和符号标注。它是迄今最大的手写乐谱数据集,也是首个包含记忆机构中真实且具代表性乐谱文档样本的数据集,适用于训练和评估端到端及基于目标检测的OMR系统,并进行性能对比。

原文摘要 · Abstract (English)

A large amount of musical heritage has been digitised by memory institutions: libraries, museums, and archives. Nevertheless, the field of Optical Music Recognition (OMR) has struggled with making this music machine-readable, despite advances in deep learning, mostly because no datasets for training systems in realistic conditions were available. The MusiCorpus dataset aims to remedy this situation by providing 1,309 pages of historical sheet music, primarily handwritten, with MusicXML transcriptions and symbol annotations. It is the largest dataset of handwritten music to date and the first dataset containing a realistic and representative sample of musical document collections from memory institutions, suitable for training and evaluating both end-to-end and object detection-based OMR systems and comparing their performance.

乐谱识别手写识别数据集文化遗产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。