解决复杂钢琴乐谱识别中的音符分声部与时间对齐难题
From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR

- 将第二阶段解码建模为结构解析问题,用概率引导搜索方法求解
- 在钢琴乐谱上实现高精度的多声部分离与小节内时序对齐
- 适合需要可编辑乐谱输出的音乐数字化项目
我们提出一种实用的两阶段光学乐谱识别(OMR)流程,重点关注第二阶段。在视觉模块生成符号与事件候选的基础上,将其解码为可编辑、可验证、可导出的乐谱结构。针对复杂复调记谱(尤其是钢琴谱),重点突破声部分离与小节内时间对齐两大瓶颈。将第二阶段解码建模为结构解码问题,采用基于拓扑识别的概率引导搜索方法(BeadSolver)作为核心算法。同时设计了一种结合程序生成与识别反馈标注的数据策略。最终构建出可用于实际OMR系统的解码组件,并为未来端到端、多模态及强化学习方法积累结构化乐谱数据。
原文摘要 · Abstract (English)
We propose a new approach for a practical two-stage Optical Music Recognition (OMR) pipeline, with a particular focus on its second stage. Given symbol and event candidates from the visual pipeline, we decode them into an editable, verifiable, and exportable score structure. We focus on complex polyphonic staff notation, especially piano scores, where voice separation and intra-measure timing are the main bottlenecks. Our approach formulates second-stage decoding as a structure decoding problem and uses topology recognition with probability-guided search (BeadSolver) as its core method. We also describe a data strategy that combines procedural generation with recognition-feedback annotations. The result is a practical decoding component for real OMR systems and a path to accumulate structured score data for future end-to-end, multimodal, and RL-style methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。