arXiv:2608.10978cs.CVcs.SD2026-08中稿 · the ISMIR 2026

首个专用于弦乐四重奏的乐谱光学识别数据集,解决多声部乐谱识别难题。

A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

论文配图:A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores
图 1 · 摘自论文原文
  • 构建首个弦乐四重奏多声部乐谱识别数据集,含24,544个系统级图像
  • 基准模型在合成数据上最低错误率3.6%,扫描数据上为5.9%
  • 提供多种编码格式与评估协议,适合多声部音乐数字化研究者

光学音乐识别(OMR)将乐谱转为数字格式。尽管单声部和钢琴谱的识别已取得显著进展,但多声部乐谱的识别仍缺乏足够研究,主要因缺少合适的数据集。本文提出面向弦乐四重奏的光学音乐识别数据集(OSSQ-OMR),基于OpenScore弦乐四重奏语料库,将数字化乐谱与IMSLP原始扫描版配对,所有图像与转录内容视觉对齐。数据集包含24,544个系统级图像和98,172个乐行级图像,源自116首弦乐四重奏作品。提供三种编码格式:扩展线性化MusicXML(LMXE)、kern和ABC。配套发布评估协议及两个代表性模型的基线结果,涵盖四个互斥测试集的随机分组。基线模型在合成输入上最低错误率为3.6%,扫描输入上为5.9%。结果表明编码方式和分割策略影响显著,基于LSTM的模型在扫描输入上的性能退化仅为基于Transformer模型的约2.6倍。

原文摘要 · Abstract (English)

Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form scores, multi-part score transcription remains underexplored, largely due to the absence of a suitable dataset. We introduce OpenScore String Quartet for Optical Music Recognition (OSSQ-OMR), the first dataset dedicated to multi-part OMR. Built on the OpenScore String Quartet corpus, OSSQ-OMR pairs digitally encoded scores with their original scanned editions from IMSLP, with all images visually aligned to their transcriptions. The dataset is released with score images at system and staff levels, and paired transcriptions in three encoding formats: Extended Linearized MusicXML (LMXE), **kern, and ABC. In total, OSSQ-OMR contains 24,544 system images and 98,172 staff images drawn from 116 string quartet scores. We accompany the dataset with a benchmark protocol and baseline results from two representative OMR models, evaluated across four random score-level splits with mutually exclusive test sets. Baselines reach OMR-NED as low as 3.6% on synthetic and 5.9% on scanned inputs; results reveal substantial effects of encoding and segmentation choices, with the LSTM-based baseline degrading on scanned inputs roughly 2.6 times less than the Transformer-based baseline.

音乐识别多声部数据集弦乐四重奏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。