arXiv:2605.24193cs.SDcs.LG2026-05被引 1

用少量标注数据激活海量无标签音乐数据,实现低监督高质量乐谱转录。

Music Transcription with (Almost) No Supervision

论文配图:Music Transcription with (Almost) No Supervision
图 1 · 摘自论文原文
  • 通过循环一致翻译框架,仅需少量配对数据作为锚点
  • 无标签音频数据带来显著性能提升,尤其在标注稀缺时
  • 训练中加入新乐器的无标签音频,可零监督提升其转录效果

高性能音乐转录模型依赖大量音符-音频配对数据,但这类数据因收集成本高、对齐困难及版权问题而稀缺。相比之下,大量未配对的音频与乐谱数据免费可得却未被利用。本文采用循环一致翻译框架,以少量配对数据作为最小锚点,释放未配对数据的全部潜力。实验发现:在监督有限的情况下,未配对数据带来显著性能提升;未配对音频的贡献大于未配对乐谱;在训练中引入某新乐器的无标签音频,可无需任何配对标注即提升该乐器的转录性能。这些结果表明,扩大未配对数据规模是解决标注稀缺乐器高质量转录的有效路径。

原文摘要 · Abstract (English)

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired audio recordings and symbolic scores are freely available but have gone unused. We adopt a cycle-consistent translation framework in which a small amount of paired data acts as a minimal anchor, unlocking the full potential of the unpaired pool. We find that: unpaired data yields surprisingly large gains, especially under limited supervision; unpaired audio contributes more than unpaired scores; incorporating unlabeled audio from a new instrument during training improves transcription for that instrument without any paired supervision. Together, these results suggest that scaling unpaired data offers a practical path toward high-quality transcription for instruments where labeled data remains scarce.

音乐转录自监督学习无监督音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。