arXiv:2605.10835cs.CVcs.LG2026-05被引 1

用合成数据训练的端到端音乐识别模型,大幅降低错误率。

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

论文配图:Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training
图 1 · 摘自论文原文
  • 通过合成数据生成与编码规范化提升训练效率
  • 仅用6小时单卡训练,59M参数模型超越千亿参数基线
  • 适合需要高精度音乐转录的研究者与开发者

光学音乐识别(OMR)将乐谱图像转为结构化文本,受限于真实扫描数据标注规模小,现有模型多依赖少样本迁移或简化的合成训练。另一挑战是编码非唯一性:在流行的Humdrum kern格式中,同一乐谱可对应多种文本表示,导致解码不确定性。本文提出Transcoda系统,采用(i)先进的合成数据生成流程,(ii)kern编码归一化以确保唯一标准形式,(iii)基于语法的解码保障输出语义正确性。该方法使一个仅5900万参数的模型在单个GPU上仅用6小时即可训练完成,性能超越数十亿参数基线。在新构建的合成乐谱基准上,其OMR-NED达18.46%,优于次优系统Legato的43.91%;在历史波兰乐谱数据集上,错误率从SMT++的80.16%降至63.97%。

原文摘要 · Abstract (English)

Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of large-scale, annotated datasets of real scans. This forces models to rely on either few-shot transfer or synthetic training pipelines that remain overly simplistic. A secondary challenge is encoding non-uniqueness: in the popular Humdrum **kern format for transcribing music, multiple different text encodings can render into the same visual sheet music. This one-to-many mapping creates a harder learning task and introduces high uncertainty during decoding. We propose Transcoda, an OMR system built on (i) an advanced synthetic data generation pipeline, (ii) a normalization of the **kern encoding to enforce a unique normal form and (iii) grammar-based decoding to ensure the syntactic correctness of the output. This approach allows us to train a compact 59M-parameter model in just 6 hours on a single GPU that outperforms billion-parameter baselines. Transcoda achieves the best score among state of the art baselines on a newly curated benchmark of synthetically rendered scores at 18.46% OMR-NED (compared to 43.91% for the next-best system, Legato) and reduces the error rate on historical Polish scans to 63.97% OMR-NED (down from 80.16% for SMT++).

音乐识别合成数据端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。