提出通用联合学习框架,同时提升音乐分离与音高估计效果。
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
- 采用两阶段训练与动态加权策略,优化双任务协同学习
- 分离效果提升0.92 SDR,音高准确率提高2.71%
- 适配多种模型架构,对数据少场景特别有效
音乐源分离与音高估计是音乐信息检索中的两个关键任务。通常音高估计依赖于源分离的输出,因此现有方法尝试联合进行这两项任务,以利用二者间的互补关系。然而,这些方法仍面临两大挑战:标注数据不足和联合学习优化困难。为此,本文提出一种模型无关的联合学习框架MAJL。该框架可适配不同模型,并包含两阶段训练与动态加权硬样本(DWHS)方法,分别应对数据稀缺与优化难题。在公开音乐数据集上的实验表明,MAJL在两项任务上均超越现有最优方法,音乐源分离的信号失真比(SDR)提升0.92,音高估计的原始音高准确率(RPA)提升2.71%。全面分析验证了各组件的有效性,并展示了框架在多种模型架构下的强泛化能力。
原文摘要 · Abstract (English)
Music source separation and pitch estimation are two vital tasks in music information retrieval. Typically, the input of pitch estimation is obtained from the output of music source separation. Therefore, existing methods have tried to perform these two tasks simultaneously, so as to leverage the mutually beneficial relationship between both tasks. However, these methods still face two critical challenges that limit the improvement of both tasks: the lack of labeled data and joint learning optimization. To address these challenges, we propose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJL is a generic framework and can use variant models for each task. It includes a two-stage training method and a dynamic weighting method named Dynamic Weights on Hard Samples (DWHS), which addresses the lack of labeled data and joint learning optimization, respectively. Experimental results on public music datasets show that MAJL outperforms state-of-the-art methods on both tasks, with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) for music source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitch estimation. Furthermore, comprehensive studies not only validate the effectiveness of each component of MAJL, but also indicate the great generality of MAJL in adapting to different model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。