arXiv:2609.05662cs.CVcs.AI2026-09

用合成数据训练手写乐谱端到端识别,关键在学布局而非模仿字形。

Full-Page Optical Music Recognition of Handwritten Monophonic Scores

论文配图:Full-Page Optical Music Recognition of Handwritten Monophonic Scores
图 1 · 摘自论文原文
  • 用生成器产出手写与印刷体乐谱,用于预训练
  • 合成数据提升性能主要靠学布局规律,非字形相似
  • 适合研究手写乐谱识别或跨域迁移的学者

全页端到端光学音乐识别旨在直接将整页乐谱转为符号记谱,避免传统流程中对五线谱分割的依赖。现有基于Transformer的模型在印刷体乐谱上表现优异,依赖大规模合成数据预训练。但其在手写乐谱上的应用仍待探索。本文研究手写单声部乐谱的全页识别,并分析合成预训练在此场景下的影响。为此,我们提出一个生成器,可生成视觉连贯的整页乐谱,支持印刷体和手写风格。在三个真实手写数据集上的实验,比较了多种全页识别流水线及不同合成预训练策略。结果表明,合成预训练的优势主要源于学习结构布局规范,而非与目标手写风格的视觉相似性。

原文摘要 · Abstract (English)

Full-page end-to-end Optical Music Recognition seeks to transcribe entire music pages directly into symbolic notation, avoiding the limitations of traditional pipelines that rely on accurate staff segmentation. Recent Transformer-based architectures have achieved strong performance on typeset scores, relying on large-scale synthetic data for pretraining. However, their applicability to handwritten music remains largely unexplored. In this work, we study full-page transcription on handwritten monophonic collections and analyze the impact of synthetic pretraining in this setting. To investigate which factors are most relevant during pretraining, we introduce a generator capable of producing visually coherent full-page scores in both typeset and handwritten styles. Experiments on three real handwritten datasets provide a comparative evaluation of several full-page pipelines and different synthetic pretraining strategies. The results suggest that the benefits of synthetic pretraining are primarily associated with learning structural layout conventions rather than with visual similarity to the target handwriting.

音乐识别手写体生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。