将乐谱图像转为带文本的符号化记谱,支持长篇自动识别。
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding

- 按阅读顺序逐系统处理乐谱,突破传统图像处理限制。
- 首次实现包含标题注释的完整符号转录,准确率超现有方法。
- 适合音乐数据挖掘、智能作曲与学术研究者使用。
我们提出一种新流水线 Legato 2,用于从乐谱图像中提取符号记谱与语义知识。Legato 2 是首个在光学乐谱识别(OMR)中采用系统级顺序处理的大规模神经模型,沿乐谱水平线条读取,而非将页面视为整体图像,从而可扩展至任意长度输入。该模型也是首个能生成包含标题、注释等嵌入文本的符号转录的 OMR 模型。流水线结合系统级分割与自回归视觉-语言模型,同时捕捉局部记谱细节与乐谱结构。在多个数据集上,Legato 2 均持续优于先前最优方法。我们还证明,符号转录可增强前沿语言模型对密集音乐文档的理解能力。Legato 2 在 OMR 及下游乐谱理解任务中均建立新基准。
原文摘要 · Abstract (English)
We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, Legato 2 consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. Legato 2 establishes new state-of-the-art performance in both OMR and downstream sheet music understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。