arXiv:2503.07352eess.AScs.LG2025-03中稿 · EUSIPCO 2025被引 1

用乐谱信息提升古典音乐分离模型在真实录音上的泛化能力

Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music

  • 用乐谱与音频谱图拼接作为输入,或仅用乐谱生成分离掩码
  • 仅用乐谱的模型在真实数据上分离效果显著优于基线方法
  • 适合追求真实场景下音乐分离性能的研究者

音乐源分离旨在将混合乐器音轨拆分为独立声部。现有模型通常仅依赖音频数据训练,但加入乐谱信息可提升分离效果。本文提出两种基于乐谱的方法:一是将乐谱与音频幅度谱图拼接作为模型输入;二是仅用乐谱计算分离掩码。模型在合成数据集SynthSOD上训练,测试于包含真实录音的URMP和Aalto anechoic orchestra数据集。结果显示,拼接输入的模型虽优于基线,但在合成到真实数据间泛化表现不佳;而仅使用乐谱的模型则在真实数据上表现出显著更好的泛化能力。

原文摘要 · Abstract (English)

Music source separation is the task of separating a mixture of instruments into constituent tracks. Music source separation models are typically trained using only audio data, although additional information can be used to improve the model's separation capability. In this paper, we propose two ways of using musical scores to aid music source separation: a score-informed model where the score is concatenated with the magnitude spectrogram of the audio mixture as the input of the model, and a model where we use only the score to calculate the separation mask. We train our models on synthetic data in the SynthSOD dataset and evaluate our methods on the URMP and Aalto anechoic orchestra datasets, comprised of real recordings. The score-informed model improves separation results compared to a baseline approach, but struggles to generalize from synthetic to real data, whereas the score-only model shows a clear improvement in synthetic-to-real generalization.

音乐分离乐谱信息泛化能力古典音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。