开源多乐器乐谱生成模型,真实音乐也能准转录。
MuScriptor: An Open Model for Multi-Instrument Music Transcription

- 合成数据预训练+真实数据微调+强化学习后处理
- 支持多种乐器组合,跨流派真实录音表现佳
- 可按需指定乐器存在性,适合音乐制作与研究
现有自动乐谱转录方法通常局限于单乐器录音,或在复杂真实音乐混音中表现不佳。尽管以往工作使用合成数据训练,但模型泛化能力差,导致真实场景下输出基本不可用。本文分析了合成数据预训练的有效性,结合真实音乐音频微调和强化学习后训练,提升模型在真实复杂混音中的表现。同时引入乐器存在性条件控制,实现定制化转录。最后发布 MuScriptor,一个开源权重的多乐器音乐转录模型,可处理跨多样音乐流派的真实世界录音。
原文摘要 · Abstract (English)
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。