提升古籍琴谱识别准确率,实现超低错误率与跨版本泛化。
KuiSCIMA v2.0: Improved Baselines, Calibration, and Cross-Notation Generalization for Historical Chinese Music Notations in Jiang Kui's Baishidaoren Gequ
- 针对稀缺不平衡数据设计字符识别模型,优化训练策略。
- 苏子谱CER降至7.1%,吕律谱仅0.9%,优于人工转录。
- 引入温度缩放校准模型,支持跨刊本泛化,适配古乐数字化研究者。
针对历史中文乐谱(如俗字谱、吕律谱)的光学乐谱识别(OMR)面临类别极度不均和训练数据有限的挑战。本文针对1202年江奎《白石道人歌曲》这一重要文献集,提出并评估了一种在稀疏不平衡数据下的字符识别模型。通过改进基线方法,将俗字谱的字符错误率(CER)从10.4%降至7.1%(面对77个高度不平衡的类别),并实现吕律谱高达0.9%的惊人低错率。模型表现超越人类转录员——人类平均CER为15.9%,最优者达7.6%。采用温度缩放实现良好校准,预期校准误差(ECE)低于0.0162。通过留一刊本交叉验证,确保在五个历史版本间具备稳健性能。同时,将KuiSCIMA数据集扩展至包含《白石道人歌曲》全部109首曲目,涵盖俗字谱、吕律谱与减字谱。研究成果推动了历史中文音乐的数字化与可及性,促进OMR领域的文化多样性,拓展至非主流音乐传统。
原文摘要 · Abstract (English)
Optical Music Recognition (OMR) for historical Chinese musical notations, such as suzipu and lülüpu, presents unique challenges due to high class imbalance and limited training data. This paper introduces significant advancements in OMR for Jiang Kui's influential collection Baishidaoren Gequ from 1202. In this work, we develop and evaluate a character recognition model for scarce imbalanced data. We improve upon previous baselines by reducing the Character Error Rate (CER) from 10.4% to 7.1% for suzipu, despite working with 77 highly imbalanced classes, and achieve a remarkable CER of 0.9% for lülüpu. Our models outperform human transcribers, with an average human CER of 15.9% and a best-case CER of 7.6%. We employ temperature scaling to achieve a well-calibrated model with an Expected Calibration Error (ECE) below 0.0162. Using a leave-one-edition-out cross-validation approach, we ensure robust performance across five historical editions. Additionally, we extend the KuiSCIMA dataset to include all 109 pieces from Baishidaoren Gequ, encompassing suzipu, lülüpu, and jianzipu notations. Our findings advance the digitization and accessibility of historical Chinese music, promoting cultural diversity in OMR and expanding its applicability to underrepresented music traditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。