arXiv:2606.09479cs.CVcs.DL2026-06中稿 · publication at the…被引 1

用合成数据提升手稿谱面识别,让旧乐谱数字化更可行

Optical Music Recognition for Real-World Manuscripts with Synthetic Data

论文配图:Optical Music Recognition for Real-World Manuscripts with Synthetic Data
图 1 · 摘自论文原文
  • 用合成图像+图结构标注训练模型,减少真实手稿标注需求
  • 在资源有限场景下,合成数据使识别准确率显著提升
  • 适合文化遗产机构做旧乐谱数字化,无需昂贵人工标注

光学乐谱识别(OMR)在模型设计上已取得显著进展,端到端方法可处理各类复杂记谱。然而,现有训练数据多为数字生成的乐谱,与图书馆等机构收藏的手稿视觉特征差异大,导致现有系统在真实场景中表现不佳。这些机构常面临资源不足,难以构建大规模领域内数据集。本文针对资源受限场景,首次建立复杂钢琴记谱手稿的识别基准。利用细粒度音乐记谱图(MuNG)标注与Smashcima合成工具,表明虽需少量真实数据作为基础,但通过合成手稿图像进行领域自适应可带来显著性能提升。关键发现是:合成符号无需与真实手稿一致,从而避免高成本细粒度标注。该工作推动了OMR向保护和传播音乐文化遗产的目标迈进。

原文摘要 · Abstract (English)

Optical Music Recognition (OMR) has seen major progress in model design, with end-to-end methods now capable of recognising notation at all levels of complexity. However, the impact of this progress has been limited by the visual domains of available training datasets, which are largely born-digital. Existing large collections of sheet music in libraries and other heritage institutions contain predominantly manuscripts, whose visual domains are highly diverse and different, so existing OMR systems fail when applied in the real world. These institutions are often resource-constrained, so large in-domain datasets cannot be expected. We provide a first baseline on real-world manuscripts with complex piano notation in the resource-constrained scenario. Using fine-grained music notation graph (MuNG) annotations and the Smashcima synthesis tool, we then show that while some direct transcriptions of in-domain data remain essential, domain adaptation using synthetic musical manuscript images brings significant improvement. Furthermore, the symbols used do not need to be in-domain, so the expensive fine-grained annotation can be avoided. We thus bring OMR closer to one of its stated goals: preserving and promoting musical cultural heritage.

光学识别手稿数字化合成数据音乐遗产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。