首个可端到端识别多页乐谱并生成人类可读乐谱的模型
LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR
- 用视觉编码器加ABC解码器的端到端架构,基于21.4万张乐谱图训练
- 在真实数据集上比之前最好方法误差降低68%(TEDn)和47.6%(OMR-NED)
- 首次支持全页/多页排版乐谱识别,输出简洁可读的ABC记谱法
我们提出Legato,一种用于光学音乐识别(OMR)的新端到端模型,任务是将乐谱图像转换为机器可读文档。Legato是首个大规模预训练的OMR模型,能够识别整页或多页排版乐谱,并首次生成简洁易读的ABC记谱法。该模型结合了预训练视觉编码器与在超过21.4万张图像数据集上训练的ABC解码器,展现出对多种排版乐谱的强大泛化能力。我们在多个数据集和指标上进行综合实验,结果表明Legato优于此前最优方法。在最接近真实的测试集上,标准指标TEDn和OMR-NED的绝对误差分别降低68%和47.6%。
原文摘要 · Abstract (English)
We propose Legato, a new end-to-end model for optical music recognition (OMR), a task of converting music score images to machine-readable documents. Legato is the first large-scale pretrained OMR model capable of recognizing full-page or multi-page typeset music scores and the first to generate documents in ABC notation, a concise, human-readable format for symbolic music. Bringing together a pretrained vision encoder with an ABC decoder trained on a dataset of more than 214K images, our model exhibits the strong ability to generalize across various typeset scores. We conduct comprehensive experiments on a range of datasets and metrics and demonstrate that Legato outperforms the previous state of the art. On our most realistic dataset, we see a 68\% and 47.6\% absolute error reduction on the standard metrics TEDn and OMR-NED, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。