arXiv:2508.13516cs.SDeess.AS2025-08中稿 · ISMIR 2025 as Late…

不用钢琴预训练,直接用小规模小提琴数据也能实现优秀乐谱转录。

Is Transfer Learning Necessary for Violin Transcription?

  • 直接在30小时小提琴数据上从零训练,不依赖钢琴预训练模型。
  • 在URMP和Bach10数据集上表现优于或媲美调优的钢琴预训练模型。
  • 证明小提琴专用数据收集与增强比跨乐器迁移更重要。

自动音乐转录(AMT)在钢琴等乐器上进展显著,主要得益于大规模高质量数据集。相比之下,小提琴的AMT因标注数据有限而研究不足。当前普遍做法是微调其他任务的预训练模型,但其在音色和演奏方式差异下的有效性尚不明确。本文探究在中等规模小提琴数据集(约30小时对齐录音,MOSA数据集)上从零训练是否能媲美微调钢琴预训练模型。采用未修改的钢琴转录架构,在URMP和Bach10数据集上实验表明,从零训练模型性能可达到甚至超过微调模型。结果表明,无需依赖钢琴预训练表示即可实现强健的小提琴自动转录,强调了针对特定乐器的数据采集与增强策略的重要性。

原文摘要 · Abstract (English)

Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.

音乐转录小提琴迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。