用专家系统识别中文简谱乐谱,自动转成可编辑音乐格式。
The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics
- 结合传统视觉分析与无监督深度学习,构建可解释的识别流程。
- 在5000首民歌上实现95.1%的音符识别准确率,歌词识别达93.1%。
- 适合音乐数字化、文化遗产保护及中文音乐数据研究者使用。
大规模光学音乐识别(OMR)研究主要聚焦西方五线谱,而中文简谱及其丰富的歌词资源长期被忽视。我们提出一个模块化专家系统流水线,将印刷版带歌词的简谱乐谱转化为机器可读的MusicXML和MIDI格式,无需大量标注训练数据。方法采用自上而下的专家系统设计,利用传统计算机视觉技术(如乐句相关性、骨架分析)发挥先验知识优势,同时集成无监督深度学习模块进行图像特征嵌入。该混合策略兼顾可解释性与精度。在《中国民歌集》数据集上评估,系统成功数字化了:(i) 超过5000首仅含旋律的歌曲(>30万音符),以及 (ii) 包含歌词的精选子集,共1400余首(>10万音符)。在旋律识别上达到音符级F1=0.951,歌词对齐识别达字符级F1=0.931。
原文摘要 · Abstract (English)
Large-scale optical music recognition (OMR) research has focused mainly on Western staff notation, leaving Chinese Jianpu (numbered notation) and its rich lyric resources underexplored. We present a modular expert-system pipeline that converts printed Jianpu scores with lyrics into machine-readable MusicXML and MIDI, without requiring massive annotated training data. Our approach adopts a top-down expert-system design, leveraging traditional computer-vision techniques (e.g., phrase correlation, skeleton analysis) to capitalize on prior knowledge, while integrating unsupervised deep-learning modules for image feature embeddings. This hybrid strategy strikes a balance between interpretability and accuracy. Evaluated on The Anthology of Chinese Folk Songs, our system massively digitizes (i) a melody-only collection of more than 5,000 songs (> 300,000 notes) and (ii) a curated subset with lyrics comprising over 1,400 songs (> 100,000 notes). The system achieves high-precision recognition on both melody (note-wise F1 = 0.951) and aligned lyrics (character-wise F1 = 0.931).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。