首个面向流行音乐的音频转乐谱数据集与模型,显著提升转写准确率。
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

- 基于预训练音频特征和数据增强,提升模型泛化能力。
- 在古典音乐上达4.98%符号错误率,流行音乐达20.92%。
- 开源数据集与代码,推动流行音乐乐谱转写研究。
现有音频转乐谱(A2S)系统多聚焦于古典音乐,流行音乐应用仍待探索。本文首次发布SheetSage-A2S数据集,包含61小时音频,对应9,468段乐谱片段,源自6,066首独特歌曲,为流行音乐A2S研究提供首个专用数据集。同时,通过引入数据增强和MuQ预训练音频特征提取模型,提升模型泛化能力与特征表达。实验显示,该模型在古典音乐的Quartets数据集上达到4.98%符号错误率(SER),显著优于现有最优方法(15.3%)。在SheetSage-A2S数据集上,流行音乐转写达到20.92% SER,成为未来研究的重要基准。相关数据集、模型与代码已公开:https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern_model。
原文摘要 · Abstract (English)
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with **kern score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3% SER from the existing state-of-the-art (Alfaro-Contreras et al. 2024). Additionally, our model achieves 20.92% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern_model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。