用主动学习减少标注量,提升古乐谱识别准确率
Experimenting active and sequential learning in a medieval music manuscript
- 基于不确定性选择样本,迭代标注并重训练模型
- 仅需少量标注数据即达全监督训练精度
- 适合古籍数字化中数据稀缺的研究者参考
光学音乐识别(OMR)是文化遗产数字化的核心,但受限于标注数据稀缺和历史手稿复杂性。本文针对中世纪音乐手稿中的目标检测与版面识别,初步研究了主动学习(AL)与顺序学习(SL)。基于YOLOv8,系统选择置信度最低的样本进行迭代标注与重训练,从单张标注图像起步,显著降低人工标注需求。实验表明,仅需极少标注即可达到全监督训练的准确率。该方法在匿名项目提供的新数据集上测试,该数据集聚焦12至16世纪意大利流行的诗歌音乐体裁‘laude’。结果显示,不确定性驱动的主动学习在当前手稿中效果不佳,提示需探索更实用的数据稀缺场景方案。
原文摘要 · Abstract (English)
Optical Music Recognition (OMR) is a cornerstone of music digitization initiatives in cultural heritage, yet it remains limited by the scarcity of annotated data and the complexity of historical manuscripts. In this paper, we present a preliminary study of Active Learning (AL) and Sequential Learning (SL) tailored for object detection and layout recognition in an old medieval music manuscript. Leveraging YOLOv8, our system selects samples with the highest uncertainty (lowest prediction confidence) for iterative labeling and retraining. Our approach starts with a single annotated image and successfully boosts performance while minimizing manual labeling. Experimental results indicate that comparable accuracy to fully supervised training can be achieved with significantly fewer labeled examples. We test the methodology as a preliminary investigation on a novel dataset offered to the community by the Anonymous project, which studies laude, a poetical-musical genre spread across Italy during the 12th-16th Century. We show that in the manuscript at-hand, uncertainty-based AL is not effective and advocates for more usable methods in data-scarcity scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。